AI music creation system based on multiple modes

By combining rule constraints with random exploration, the multimodal AI music creation system dynamically adjusts the controllability and unpredictability of generated music, solving the problems of formulaic and logically chaotic music generation in existing technologies. This achieves improvements in artistry and user satisfaction, while also supporting complex music styles and copyright protection.

CN121617367APending Publication Date: 2026-03-06XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511925838.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing music generation algorithms struggle to strike a balance between algorithmic controllability and artistic unpredictability when generating music, resulting in formulaic and unoriginal musical works, or logically confused works that conflict with cultural traditions.

Method used

The system employs a multimodal AI music creation system, including a user control layer, a controllability engine, an unexpectedness engine, a dynamic balancer, and a music generation core. By combining rule constraints with random exploration, it dynamically adjusts the controllability and unexpectedness of generated music. It also introduces a multi-dimensional balance adjustment interface and a graphical user interface, allowing users to adjust the generation strategy in real time.

Benefits of technology

It enhances the artistry and user satisfaction of generated music works, avoids mechanization and monotony, supports the generation of complex music styles, and ensures originality through copyright protection mechanisms, significantly improving creation efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121617367A_ABST
    Figure CN121617367A_ABST
Patent Text Reader

Abstract

The invention discloses an AI music creation system based on multiple modes, and the system comprises a user control layer which is used for receiving the input of a user, and transmitting a control signal to a controllability engine and an accidental engine; and the controllability engine is connected with a rule constraint domain, and the rule constraint domain comprises a music theory rule base and a culture compatible filter and is used for performing rule constraint on music generation based on the music theory rule base and the culture compatible filter. According to the AI music creation system based on multiple modes, through a closed-loop feedback mechanism of the artistry evaluation module and the strategy adjustment module, a music generation strategy can be dynamically adjusted in real time, the mechanization and monotonicity problems of music generation are effectively avoided, the artistry of music works and the user satisfaction degree are remarkably improved, and then the user experience is improved. According to the method, the dimension collaborative optimizer is introduced, consistency analysis is carried out on music dimensions such as chord proceeding, melody development and rhythm design, dimension conflicts are effectively prevented, and the generated music works are more harmonious in structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of music information processing technology, specifically to an AI music creation system based on multimodality. Background Technology

[0002] With the widespread application of artificial intelligence technology in the field of music creation, algorithm-based music generation systems have gradually become a research hotspot. Existing music generation algorithms mostly focus on generating musical works with certain patterns through preset rules or machine learning models, aiming to ensure the rationality and standardization of the generated music in terms of structure, harmony, etc. For example, deep learning models are used to learn patterns in a large amount of music data, thereby generating works that conform to common music theories.

[0003] However, the aforementioned existing technologies have obvious drawbacks. On the one hand, their over-reliance on fixed rules or existing patterns results in generated musical works that exhibit a high degree of formulaic characteristics, lacking innovation and artistic surprise, and failing to meet users' demands for personalized and novel musical works; On the other hand, some algorithms that attempt to introduce random elements to increase unpredictability often lack effective control mechanisms, leading to logical inconsistencies and conflicts with cultural traditions in the generated musical works, thus failing to achieve artistic breakthroughs while ensuring musical quality. Therefore, finding a balance between algorithmic controllability and artistic unpredictability has become a pressing technical challenge. Summary of the Invention

[0004] Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a multimodal AI music creation system that has advantages such as rule-based music generation, a balance between algorithmic controllability and artistic unpredictability, and solves the problems of logical inconsistencies and conflicts with cultural traditions in musical works.

[0005] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: a multimodal AI music creation system, comprising: S1. User Control Layer: Used to receive user input and transmit control signals to the controllability engine and the contingency engine respectively; S2. Controllability Engine: Connects the rule constraint domain, which includes a music theory rule base and a cultural compatibility filter, and is used to impose rule constraints on music generation based on the music theory rule base and the cultural compatibility filter; S3. Unexpectedness Engine: Connects to a random exploration domain, which includes a chaotic mathematical perturbation module and a style confusion mutation module, used to introduce unexpected elements through the chaotic mathematical perturbation module and the style confusion mutation module; S4. Dynamic Equalizer: Connects the rule constraint domain, the random exploration domain and the music generation core. The dynamic equalizer includes an entropy feedback loop and an artistic evaluation unit, which are used to achieve dynamic balance control of controllability and unpredictability, and input the balance signal into the music generation core. S5. Music Generation Core: Used to generate music based on the signal input from the dynamic equalizer; S6. Dual-engine collaborative control mechanism: Set controllability and unpredictability strategies for dimensions such as chords, melody, and rhythm, and achieve balance through routine index threshold, ornament density slider, and cultural compliance detection.

[0006] Preferably, the dual-engine collaborative control mechanism sets controllability strategies, unexpected strategies, and balance rules for the dimensions of chords, melody, and rhythm, respectively. The balance strategy includes routine index threshold triggering, ornament density slider adjustment, and cultural compliance detection.

[0007] Preferably, the calculation of music entropy value is used to measure unexpectedness, obtain the safety boundary set by the user, make decisions through an artistic evaluation model, and break the constraint when the artistic evaluation result is higher than the set threshold. Entropy feedback control is performed based on music entropy and a safety boundary. When the music entropy exceeds the safety boundary by a certain multiple, rule constraints are activated; otherwise, the unpredictability is enhanced. The core algorithm of the dynamic balancer is: in, The entropy value of music is represented by the unexpectedness calculated through the probability distribution of musical notes. Indicates the safety boundary threshold. Represents the results of the artistic assessment. This represents the threshold for artistic evaluation.

[0008] Preferably, it also includes a three-tiered artistic surprise injection model, which includes a subtle decorative surprise level, a local structural surprise level, and a disruptive aesthetic breakthrough level. Each level corresponds to different surprise types, technical implementations, and controllable parameters. The mathematical expression of the three-tiered artistic surprise injection model is: in, These represent unexpected parameters at the levels of subtle decoration, local structure, and disruptive breakthrough.

[0009] Preferably, it also includes a culturally sensitive unpredictability constraint system, which is used to impose unpredictability constraints on music with different cultural characteristics.

[0010] Preferably, it includes an art value assessment feedback loop, which is used to assess the artistry of the generated music in real time and adjust constraints and generation strategies based on the assessment results.

[0011] Preferably, it also includes a phased deployment strategy and a dual mechanism for copyright protection. The phased deployment strategy is used to release information about unexpected events in stages, and the dual mechanism for copyright protection includes tracing the source of the work and holding those who violate copyright accountable. The mathematical expression for the phased deployment strategy and the dual mechanism for copyright protection is as follows: in, Indicates the deployment time or stage progress. Stages are used to identify different deployment stages or process states of a music creation system. This represents the threshold, a critical value used to determine whether to enable unexpectedness. It is triggered when the deployment time or phase progress T exceeds the threshold. Enable unexpectedness when it is enabled; otherwise disable unexpectedness.

[0012] Preferably, the multi-dimensional balance adjustment interface is used to provide a graphical user interface, allowing users to independently set the controllability level and unexpectedness level for chord progression, melody development, and rhythm design dimensions, and to adjust the routine index threshold and ornament density in the balance strategy in real time through sliders or input boxes.

[0013] Preferably, it also includes an inter-dimensional co-optimizer, which is used to analyze the musical consistency between chord progression, melody development, and rhythm design dimensions, and dynamically adjust the controllability engine output and unexpectedness engine injection of each dimension based on the balance strategy to prevent dimensional conflicts.

[0014] (III) Beneficial Effects Compared with existing technologies, this invention provides a multimodal AI music creation system with the following advantages: 1. This multimodal AI music creation system, through a closed-loop feedback mechanism of an artistic evaluation module and a strategy adjustment module, can dynamically adjust the music generation strategy in real time, effectively avoiding the mechanical and monotonous problems of generated music, and significantly improving the artistry and user satisfaction of musical works. Secondly, this invention introduces a dimensional collaborative optimizer to perform consistency analysis on musical dimensions such as chord progressions, melody development, and rhythm design. By dynamically adjusting the controllability engine output and unexpectedness engine injection of each dimension, it effectively prevents dimensional conflicts, making the generated musical works more structurally harmonious and richer in musical expression. Simultaneously, this invention supports the generation of complex musical styles, meeting the needs of professional music creation. Furthermore, through an efficient copyright protection mechanism, it ensures the traceability of the originality of musical works and provides a basis for accountability in the event of unauthorized use, providing users with comprehensive copyright protection.

[0015] 2. This multimodal AI music creation system allows users to adjust parameters such as "formula index" and "ornament density" in real time through a multi-dimensional balance adjustment interface and a graphical user interface, achieving a high degree of personalized customization of the music generation process and significantly improving the system's usability and user experience. Simultaneously, through an artistic evaluation feedback loop and strategy adjustment mechanism, this invention can rapidly iterate and optimize the generated results, significantly reducing the number of manual adjustments and trial-and-error attempts required by users during the music creation process, and greatly improving creation efficiency. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the overall architecture of an AI music creation system based on multimodality, as proposed in this invention. Figure 2 This is a flowchart of the main music generation process of an AI music creation system based on multimodality proposed in this invention. Figure 3 This is a flowchart illustrating the artistic evaluation and strategy adjustment feedback process of a multimodal AI music creation system proposed in this invention. Figure 4 This is a flowchart illustrating the multi-dimensional balance adjustment and collaborative optimization process of a multimodal AI music creation system proposed in this invention. Figure 5 This is a flowchart of the phased deployment strategy and second embodiment of the multimodal AI music creation system proposed in this invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1 This invention provides a technical solution: a multimodal AI music creation system, comprising the following modules: User input module: Based on users' historical input data and preference analysis, it uses collaborative filtering algorithms to recommend suitable templates. The template design employs a graphical interface, allowing users to quickly select parameters using components such as sliders and drop-down menus. The slider uses a high-precision linear interpolation algorithm to ensure the fineness of parameter adjustment. A ResNet-50 model is used to extract color, texture, and emotional features from the image. These visual features are then converted into musical tonality and rhythm parameters through a feature mapping matrix. Based on the image's color and warm / cool tones, corresponding musical compositions are generated. A BERT model is used to analyze text semantics, and combined with an emotional dictionary, musical emotional vectors are generated to guide the design of the melody.

[0019] Music generation module: This module is based on the TensorFlow or PyTorch deep learning framework and adopts an architecture that combines variational autoencoder (VAE) and generative adversarial network (GAN). VAE is used to learn the latent spatial distribution features of music, and GAN is used to generate high-quality music sequences. An attention mechanism is introduced to optimize key user parameters, and transfer learning is used to achieve rapid fine-tuning of specific music styles.

[0020] The artistic evaluation module uses convolutional neural networks (CNNs) to extract low-level features from the music spectrogram, analyzes the time-series features of the music through recurrent neural networks (RNNs) and their variants, integrates an emotion computing model to evaluate the emotional intensity of the music, applies a complexity calculation algorithm to measure the compositional complexity of the music, and uses hashing algorithms and similarity calculation methods to evaluate the originality of the music. The evaluation results are output in the form of multi-dimensional vectors to evaluate the artistry of the musical works, and the evaluation results are fed back to the strategy adjustment module. Emotional matching degree: calculated by the cosine similarity between the melody fluctuations and the text's emotional vector; Structural innovation: Based on the LSTM model, the melody direction is predicted, and the entropy of the difference between the actual generated result and the predicted result is calculated; Cultural compliance: Compare the matching degree of the generated music's mode and rhythmic pattern with the target cultural database; Algorithm Model: Employs a multi-layer perceptron to fuse multi-dimensional indicators, outputting an artistic score of 0-10, with a threshold value... The default score is 7.

[0021] Strategy Adjustment Module: Based on the artistic evaluation results and preset rules, adjust the parameters and strategies of the music generation module, specifically as follows: The music generation process is modeled as a Markov decision process; Using artistic evaluation results as a reward signal; Policy optimization is achieved by combining the Deep Q-Network (DQN) algorithm with an expert rule base.

[0022] Dimensional Co-optimizer: Used to analyze the consistency between chord progressions, melody development, and rhythm design dimensions, and dynamically adjusts the controllability engine output and unexpected engine injection of each dimension through a balancing strategy to prevent dimensional conflicts. Specifically: Abstractly define constraints in dimensions such as chords, melodies, and rhythms. Adjust the output of the controllability and unpredictability engine using a backtracking search algorithm. An adaptive weighting algorithm is applied to achieve multi-dimensional balance optimization.

[0023] Phased Deployment Strategy Module: This module is used to open up the degree of unexpectedness in stages, divide the music generation into different stages, switch states according to stage characteristics and preset conditions, and dynamically adjust the degree of unexpectedness in music creation.

[0024] The copyright protection mechanism includes source tracing and accountability for unauthorized use. Blockchain technology records the hash value of music metadata, and digital watermarking technology uses a spread spectrum algorithm to embed copyright information.

[0025] Multi-dimensional balance adjustment interface: This module implements a graphical user interface based on WebGL: It provides a graphical user interface that allows users to independently set the controllability level and unexpectedness level for chord progression, melody development, and rhythm design dimensions, realizes real-time parameter rendering and slider control, and adjusts the routine index value and ornament density in the balance strategy in real time through sliders or input boxes, and realizes front-end and back-end communication through WebSocket.

[0026] Implementation steps: User input: Users input their music requests through the user interface; Controllable parameters: Users can adjust parameters such as pitch, scale, and rhythm to control the overall style of the music and generate corresponding works based on the colors and warm / cool tones of the image. Unexpected guidance: Users can choose different unexpected guidance strategies, such as introducing random notes or chords, to increase the fun of the musical works; Animation generation: Animation element type, using graphics libraries (such as OpenGL, Pygame) or deep learning models (such as GAN) to generate animation elements, and dynamically adjust animation parameters (such as color, size, motion trajectory) based on music characteristics. Interactive user controls: Animation style (e.g., abstract, realistic), visual effects (e.g., color saturation, blur effect), dynamic parameters (e.g., animation speed, graphic size). Use graphical interface frameworks (e.g., Qt, Tkinter) to design the user interface, provide controls such as sliders and buttons, and allow users to adjust animation parameters (e.g., style, color saturation). Multimodal data fusion: mapping relationship between musical and visual features, collaborative generation mechanism, and real-time adjustment of animation element parameters (such as color, size, and motion trajectory) based on changes in musical features. For example, the visual impact of a musical climax can be enhanced by increasing the number and speed of animation elements. The specific steps of data fusion: Feature extraction: Extracting rhythmic, melodic, and chordal features from music; Feature mapping: Mapping musical features to visual features (such as color, shape, motion trajectory); Collaborative generation: Combines mapping relationships to generate dynamic animations and adjusts animation elements in real time to match changes in music, thereby building an optimized architecture for animation generation and achieving animation generation effects for different music genres (such as classical music and pop music). Verify whether animated elements match musical characteristics. Use parallel computing techniques (such as multithreading and GPU acceleration) to improve animation generation speed.

[0027] Optimize the motion trajectory algorithm for animated elements to reduce computational overhead.

[0028] Artistic Evaluation Value: Users can input their expectations for the artistic quality of a musical work, such as emotional intensity and complexity; Music generation: The music generation module adjusts the module's parameters and strategies based on user input and strategies to generate musical works; Artistic Assessment: The artistic assessment module evaluates the artistry of musical works based on preset assessment criteria and models. Assessment criteria may include emotional intensity, complexity, and originality. The artistic assessment module then feeds back the assessment results to the strategy adjustment module.

[0029] Strategy Adjustment: The strategy adjustment module adjusts the parameters and strategies of the music generation module based on the artistic evaluation results and preset rules. Adjustment rules may include: If the artistic evaluation score is lower than the user's expectations, the level of unexpected guidance will be increased to enhance the fun of the musical work.

[0030] If the artistic evaluation value is too high, reduce the degree of unexpected guidance to avoid making the musical work too chaotic.

[0031] Adjust the controllability and unpredictability of each dimension based on the user's preferences for different dimensions, such as melody, chords, and rhythm.

[0032] Artistic value assessment feedback loop: used to evaluate the artistry of generated music in real time, and adjust constraints and generation strategies based on the assessment results.

[0033] Phased deployment strategy: in, Indicates the deployment time or stage progress. Stages are used to identify different deployment stages or process states of a music creation system. This represents the threshold, a critical value used to determine whether to enable unexpectedness. It is triggered when the deployment time or phase progress T exceeds the threshold. Enable unexpectedness when it is enabled; otherwise disable unexpectedness.

[0034] Multi-dimensional balance adjustment: Users can independently set the controllability level and unexpectedness level for melody development and rhythm design through the graphical user interface, and adjust the routine index value and ornament density in the balance strategy in real time through sliders or input boxes.

[0035] Dimensional Co-optimization: The dimensional co-optimizer is used to analyze the consistency between chord progression, melody development, and rhythm design dimensions, and dynamically adjusts the controllability engine output and unexpectedness engine injection of each dimension through the balancing strategy, thereby preventing dimensional conflicts.

[0036] Music Output: The music generation module outputs the generated music to the user. Users can then edit and adjust the music to meet their specific needs.

[0037] Example 2: Suppose a user wants to generate a romantic piano piece with high emotional intensity. The user can input the following parameters through the user interface: Step 1: User Input Processing: Provide intelligent recommendation function settings parameter explanation and prompts based on collaborative filtering algorithm to assist users in setting parameters; Step 2, Adjustment of controllable parameters: A hierarchical control architecture is adopted, and basic parameters are converted through a mapping table. Rhythm parameters are selected in combination with the database, and parameter compatibility constraints are checked. Step 3, Unexpectedness Guidance: Combining random sampling and rule constraints, determine the random proportion based on the degree of unexpectedness, and apply the Markov chain model to ensure the rationality of elements; Step 4: Input of artistic evaluation value: Supports fuzzy evaluation function, adopts semantic fuzzy matching technology, and provides sample music to help set the expected value; Step 5, Music Generation Process: Parallel computing technology is introduced to decompose the task, supporting multi-core or distributed execution, and dimensional weight combination is achieved through fusion algorithm.

[0038] Step Six: Artistic Evaluation: Implement real-time evaluation functionality, generate evaluation curves, and support incremental learning to update the evaluation model. Step 7: Strategy Adjustment: Record the adjustment history and form a strategy library through meta-learning to support manual intervention by users; Step 8: Artistic Value Assessment Feedback Closed Loop: Achieve adaptive adjustment of the feedback cycle and introduce predictive models to optimize the adjustment strategy; Step 9, Phased Deployment Strategy: Set multiple trigger conditions to achieve gradual transition adjustments and provide phased goal prompts; Step 10, Multi-dimensional Balance Adjustment: Provides real-time parameter preview and comparison functions, and supports saving and selecting solutions; Step 11, Dimensional Collaborative Optimization: Achieve visual monitoring of dimensional relationships, and provide conflict highlighting and adjustment suggestions; Step 12, Music Output: Configure the format conversion engine to support multiple output formats, provide quality adjustment functions, and output with copyright information files.

[0039] The music generation module generates a piano piece based on user input and the parameters and strategies of the strategy adjustment module. The artistry assessment module evaluates the emotional intensity of the piano piece and feeds the assessment results back to the strategy adjustment module. If the assessment results show that the emotional intensity is lower than the user's expectation, the strategy adjustment module increases the proportion of dissonant notes to increase the emotional intensity of the musical work. After multiple iterations, the final generated piano piece meets the user's emotional intensity requirements.

[0040] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0041] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-modal based AI music composition system, characterized in that, Comprise: S1. User control layer: for receiving user input and transmitting control signals to controllability engine and unexpectedness engine respectively; S2. Controllability engine: connected to rule constraint domain, which includes music theory rule base and cultural compatibility filter, for rule constraint on music generation based on music theory rule base and cultural compatibility filter; S3. Unexpectedness engine: connected to random exploration domain, which includes chaotic mathematical disturbance module and style confusion variation module, for introducing unexpected elements through chaotic mathematical disturbance module and style confusion variation module; S4. Dynamic balancer: connected to rule constraint domain, random exploration domain and music generation core, which includes entropy value feedback loop and artistic evaluator, for realizing dynamic balance regulation of controllability and unexpectedness, and inputting balance signal to music generation core; S5. Music generation core: for generating music according to signal input by dynamic balancer; S6. Double-engine collaborative control mechanism: for setting controllability and unexpectedness strategies for chord, melody and rhythm dimensions, and realizing balance through routine index threshold, ornament density slider and cultural compliance detection.

2. The multi-modal based AI music composition system of claim 1, wherein: The double-engine collaborative control mechanism sets controllability strategy, unexpectedness strategy and balance rule for chord, melody and rhythm dimensions respectively, and the balance strategy includes routine index threshold triggering, ornament density slider adjustment and cultural compliance detection.

3. The multi-modal based AI music composition system of claim 1, wherein: The calculated music entropy is used to measure unexpectedness, obtain user-set safety boundary, make decision through artistic evaluation model, and break through constraint when artistic evaluation result is higher than set threshold; Based on the music entropy value and the safety boundary, the entropy value feedback regulation is carried out. When the music entropy value is greater than the safety boundary by a certain multiple, the rule constraint is enabled, otherwise the unexpectedness is enhanced. The core algorithm of the dynamic balancer is: where H represents a music entropy value, which is calculated by a note probability distribution to represent unexpectedness, and S represents a safety boundary threshold value, represents an artistic evaluation result, represents an artistic evaluation threshold value.

4. The multi-modal based AI music composition system of claim 1, wherein: The three-level artistic unexpectedness injection model comprises a fine decoration unexpectedness level, a local structure unexpectedness level and a subversive aesthetic breakthrough level, each level corresponding to different unexpectedness types, technical implementation and controllable parameters; the mathematical expression of the three-level artistic unexpectedness injection model is: wherein, respectively represent the unexpectedness parameter of the fine decoration, the local structure, and the disruptive breakthrough level.

5. The multi-modal based AI music composition system of claim 1, wherein: It also includes a culture-sensitive unexpectedness constraint system for unexpectedness constraint on music with different cultural characteristics.

6. The multi-modal based AI music composition system of claim 1, wherein: It also includes an artistic value evaluation feedback loop for real-time evaluation of artistic of generated music and adjustment of constraint and generation strategy based on evaluation result.

7. The multi-modal based AI music composition system of claim 1, wherein: Also include stage deployment strategy and copyright protection double mechanism, the stage deployment strategy is used for opening the degree of unexpectedness in stages, the copyright protection double mechanism includes work tracing and overreach accountability, the mathematical expression of the stage deployment strategy and the copyright protection double mechanism is: wherein, represents a deployment time or a stage progress, represents a stage, for identifying different deployment stages or process states of the music creation system, represents a threshold, which is a critical value for determining whether to enable the unexpectedness, when the deployment time or the stage progress T exceeds the threshold , the unexpectedness is enabled; otherwise, the unexpectedness is disabled.

8. The multi-modal based AI music composition system of claim 1, wherein: The multi-dimensional balance adjustment interface is used to provide graphical user interface, allowing user to independently set controllability level and unexpectedness level for chord progression, melody development and rhythm design dimensions, and adjust routine index threshold and ornament density in balance strategy through slider or input box in real time; at the same time, the graphical user interface also supports user to real-time adjust animation style, visual effect and dynamic parameters of MV animation module.

9. The multi-modal based AI music composition system of claim 1, wherein: It also includes an inter-dimensional collaborative optimizer for analyzing music consistency between chord progression, melody development and rhythm design dimensions, and dynamically adjusting controllability engine output and unexpectedness engine injection of each dimension based on balance strategy to prevent dimension conflict.

10. A music composition system characterized by comprising: It also includes MV animation module for generating dynamic visual elements based on music characteristics and working with music creation system, specifically including: Music feature extraction unit for extracting rhythm, melody and chord progression features of music; Animation generation unit for generating dynamic visual elements based on music characteristics, including but not limited to graphics, color and motion trajectory; An interactive control unit is configured to allow a user to adjust animation styles, visual effects and dynamic parameters through a graphical interface; A multi-modal fusion unit is configured to combine music features and visual features to generate a coordinated and consistent music and animation work.