The application provides a method and
system for generating personalized music based on AI, and related equipment, the method comprising receiving a multi-
modal creation request of a user, the request including one or more of a text description, an image, a video, and a humming melody; preprocessing and cross-
modal semantic analysis of the multi-
modal creation request to generate structured music semantic parameters; then using a
structure generation model to generate a symbolic intermediate representation of information defining the structure of the musical form, harmony, and melody based on the music semantic parameters, and then rendering a multi-track audio containing specific instrument
timbre and performance details through a
texture rendering model conditioned on the intermediate representation; finally, adaptive mixing and mastering of the audio to output a final music file. The method of the application realizes diversified
user input, controllable music structure, and automatic post-
processing, thereby improving the intelligent level of personalized music production.