Music file generation method and device, electronic equipment and storage medium

By extracting the music feature features in the modal data and mapping them into music representation sequences, the problem that existing AI music generation technology is difficult to accurately express user needs is solved, and more efficient and fine-grained music file generation is achieved.

CN119943008APending Publication Date: 2025-05-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311477443.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing AI music generation technology is difficult to accurately express user needs, and the generated music is difficult to conform to the user's emotions and music style.

Method used

By obtaining modal data, extracting the characteristics of music elements, mapping them into music representation sequences, and decoding them to generate music files. This method supports multimodal data input, improving the accuracy and fine-grainedness of music file generation.

Benefits of technology

It improves the accuracy and freedom of music file generation, enhances the degree of matching between input modal data and output music file, and ensures that the generated music is more in line with user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943008A_ABST
    Figure CN119943008A_ABST
Patent Text Reader

Abstract

The invention provides a music file generation method and device, electronic equipment and a storage medium. The method comprises the steps that at least one type of modal data is obtained, and each mode is associated with at least one music element; performing feature extraction processing on the at least one modal data to obtain at least one modal feature; according to the type of the music element associated with each modal feature, performing mapping processing on each modal feature to obtain an element feature of each music element associated with the modal feature; each element feature is mapped into a music representation sequence, and the music representation sequence comprises sound information corresponding to different moments in the music file; and decoding the music representation sequence to obtain a music file. According to the invention, the accuracy of generating the music file can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to computer technology, and in particular to a method, device, electronic device and storage medium for generating music files. Background Art

[0002] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence can be used to generate text, images or music.

[0003] In the related art, AI music generation technology can generate a variety of music, but the music generated without control is difficult to meet the needs of users, and the music generated based on input data is difficult to express the emotions required by users, affecting the user experience. In the related art, there is currently no artificial intelligence music generation solution that can accurately express user needs. Summary of the invention

[0004] The embodiments of the present application provide a method, apparatus, device, computer-readable storage medium, and computer program product for generating a music file, which can improve the accuracy of generating a music file.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present application provides a method for generating a music file, the method comprising:

[0007] Acquire at least one modality data, wherein each modality is associated with at least one music element;

[0008] Performing feature extraction processing on the at least one modal data to obtain at least one modal feature;

[0009] According to the type of the music element associated with each modal feature, mapping processing is performed on each modal feature to obtain element features of each music element associated with the modal feature;

[0010] Mapping each of the element features into a music representation sequence, wherein the music representation sequence includes sound information corresponding to different moments in the music file;

[0011] The music representation sequence is decoded to obtain a music file.

[0012] The present application provides a device for generating a music file, including:

[0013] A data acquisition module configured to acquire at least one modality of data, wherein each modality is associated with at least one musical element;

[0014] A feature extraction module, configured to perform feature extraction processing on the at least one modal data to obtain at least one modal feature;

[0015] The feature extraction module is further configured to perform mapping processing on each of the modal features according to the type of the music element associated with each of the modal features, so as to obtain element features of each music element associated with the modal feature;

[0016] A decoding module configured to map each of the element features into a music representation sequence, wherein the music representation sequence includes sound information corresponding to different moments in the music file;

[0017] The decoding module is further configured to decode the music representation sequence to obtain a music file.

[0018] An embodiment of the present application provides an electronic device, the electronic device comprising:

[0019] A memory for storing computer executable instructions;

[0020] The processor is used to implement the method for generating a music file provided in an embodiment of the present application when executing the computer executable instructions stored in the memory.

[0021] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for implementing the method for generating a music file provided in the embodiment of the present application when executed by a processor.

[0022] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the method for generating a music file provided in the embodiment of the present application is implemented.

[0023] The embodiments of the present application have the following beneficial effects:

[0024] The modal data is mapped to the feature of the music element, the feature of the music element is mapped to the music representation sequence, and the music file is generated based on the music sequence. Multimodal controllable, supports the generation of music files through different modal data, compared with the solution of generating music files based on single modal data in related technologies, it improves the freedom of obtaining music files; modal data is mapped to music elements, and then the music file is determined by the music elements, which improves the granularity of the music file generated by modal data. Compared with the solution of related technologies that rely on music materials, it saves computing resources and improves the matching degree between the input modal data and the output music file, so that the generated music file can be more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a schematic diagram of an application mode of the method for generating a music file provided in an embodiment of the present application;

[0026] Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0027] FIG. 3A to FIG. 3F is a flowchart of a method for generating a music file provided in an embodiment of the present application;

[0028] Figure 4 is a structural diagram of a music file model provided in an embodiment of the present application;

[0029] Figure 5 This is an optional flowchart of the method for generating a music file provided in an embodiment of the present application;

[0030] Figure 6 It is a schematic diagram of the structure of the parser provided in the embodiment of the present application;

[0031] Figure 7 is a schematic diagram of the structure of a music representation sequence provided in an embodiment of the present application;

[0032] Figure 8 It is a schematic diagram of the principle of the method for generating a music file provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0034] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0035] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0036] It should be pointed out that the relevant data collection and processing (for example, modal data uploaded by users) in the embodiments of the present application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in instances, and the informed consent or separate consent of the personal information subject should be obtained, and subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0038] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0039] 1) Modality: Each source or form of information can be called a modality. For example, the media of information include voice, image, text, etc.; the sources of information include various sensors, such as radar, infrared, accelerometer, etc. Each of the above can be called a modality.

[0040] 2) Convolutional Neural Networks (CNN) is a type of feed forward neural network (FNN) that includes convolution calculations and has a deep structure. It is one of the representative algorithms of deep learning. Convolutional neural networks have the ability of representation learning and can perform shift-invariant classification on input images according to their hierarchical structure.

[0041] 3) Transformer model: The Transformer model is a deep neural network model based on the self-attention mechanism. It is widely used in various tasks in the field of natural language processing, such as text classification, machine translation, and question-answering systems. The model can convert input sequences into output sequences while retaining important information in the input sequence. Since the Transformer model performs well in processing long texts, it has been widely used in the field of Chinese natural language processing. Compared with traditional recurrent neural networks (RNNs) and convolutional neural networks (CNNs), the Transformer model can be calculated in parallel to speed up training. It is widely used in tasks such as natural language processing, speech recognition, and image generation.

[0042] 4) Musical elements: basic musical elements refer to the various elements that constitute music, including the pitch, duration, strength and timbre of the sound. These basic elements are combined to form the commonly used "formal elements" of music, such as rhythm, melody, harmony, dynamics, speed, mode, form, texture, emotion, style, rhythm and notes. The musical elements in the embodiments of the present application are specifically formal elements.

[0043] 5) Emotion is a general term for a series of subjective cognitive experiences. It is a person's attitude towards objective things and the corresponding behavioral response. It is generally believed that emotion is a psychological activity mediated by individual wishes and needs. The types of emotions expressed by music include: joy, laziness, sadness, playfulness, excitement, romance, tranquility, horror, grandeur, etc.

[0044] 6) Musical style, or music style, music genre, music type, music school, etc., refers to the unique and representative appearance of a musical work as a whole, i.e. music type. For example: pop, classical, Chinese style, jazz, heavy metal, rock, light music.

[0045] 7) Rhythm refers to the format of level and oblique tones and the rhyme rules in poetry, which is extended to the rhythmic rules of sound in the field of music.

[0046] 8) Notes. Notes are symbols used to record the progression of sounds of different lengths. Whole notes, half notes, quarter notes, eighth notes, and sixteenth notes are the most common notes. They are the most important elements in the five-line staff.

[0047] 9) Musical Instrument Digital Interface (MIDI) format, a set of instructions for recording sound information. Based on the sound information carried in the instructions, the sound card can reproduce music.

[0048] 10) One-Hot Encoding refers to using N bits of 0 or 1 to encode N states. Each state has its own independent representation, and only one bit is 1, and the other bits are 0.

[0049] The embodiments of the present application provide a method for generating a music file, a device for generating a music file, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of generating a music file.

[0050] The following describes an exemplary application of an electronic device provided by an embodiment of the present application. The electronic device provided by an embodiment of the present application can implement a terminal device, such as a laptop, a tablet computer, a desktop computer, a set-top box, a smart TV, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), a vehicle terminal, a virtual reality (VR) device, an augmented reality (AR) device, and other types of user terminals, and can also be implemented as a server. Below, an exemplary application when the electronic device is implemented as a terminal device or a server will be described.

[0051] refer to Figure 1 , Figure 1 Schematic diagram of the application mode of the method for generating a music file provided in an embodiment of the present application; for example, Figure 1 The server 200, network 300, terminal device 400 and database 500 are involved. The terminal device 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0052] For example, the modal data may be at least one of an image, text, a tag value, an audio file, or a video. The server 200 may be a server of a music creation platform or a short video platform, and the database 500 stores a large amount of music data or music files produced by the server 200.

[0053] In some embodiments, a user uploads modal data to a server 200 via a terminal device 400. The server 200 calls a method for generating a music file provided in an embodiment of the present application based on the modal data, generates a corresponding music file, and sends the music file to the terminal device.

[0054] In some embodiments, the method for generating a music file in the embodiment of the present application can also be applied in the following application scenarios: 1. Video soundtrack, for example, when a film or TV series clip or a short video has no soundtrack, the user can call the method for generating a music file provided in the embodiment of the present application through a terminal device to generate a music file that matches the video content as the soundtrack of the video. 2. Music creation, for example, a user can record humming as an audio file through a terminal device, and based on the audio file, the terminal device calls the method for generating a music file provided in the embodiment of the present application to generate a new music file, so that users without music foundation can also create music.

[0055] The embodiment of the present application can be implemented through blockchain technology. The music files generated by the embodiment of the present application can be uploaded to the blockchain for storage, and the reliability of the music files can be guaranteed by the consensus algorithm. Blockchain is a new application mode of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods, each of which contains information about a batch of music files, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.

[0056] The embodiments of the present application can be implemented through database technology. In short, a database can be regarded as an electronic file cabinet where electronic files are stored. Users can add, query, update, delete, etc. data in the files. The so-called "database" is a collection of data that is stored together in a certain way, can be shared with multiple users, has as little redundancy as possible, and is independent of the application program.

[0057] A database management system (DBMS) is a computer software system designed for managing databases. It generally has basic functions such as storage, retrieval, security, and backup. Database management systems can be classified according to the database model they support, such as relational, XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters, mobile phones; or according to the query language used, such as Structured Query Language (SQL), XQuery; or according to performance focus, such as maximum scale, maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMS can cross categories, for example, supporting multiple query languages ​​at the same time.

[0058] The embodiments of the present application can also be implemented through cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model application. It can form a resource pool, which is used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, as well as the promotion of search services, social networks, mobile commerce and open collaboration, each item may have its own hash code identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately, and all kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0059] In some embodiments, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The electronic device may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal device and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.

[0060] See also Figure 2 , Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may be Figure 1 The server 200 or the terminal device 400 of the present application is described by taking the terminal device 400 as an example. Figure 2 The terminal device 400 shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .

[0061] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0062] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0063] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0064] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0065] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0066] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0067] A network communication module 452, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.;

[0068] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., display screen, speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripherals and displaying content and information);

[0069] The input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more input devices 432 and translate the detected inputs or interactions.

[0070] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2 A music file generating device 455 stored in a memory 450 is shown, which may be software in the form of a program or plug-in, including the following software modules: a data acquisition module 4551, a feature extraction module 4552, and a decoding module 4553. These modules are logical and can therefore be arbitrarily combined or further split according to the functions implemented.

[0071] In some embodiments, the terminal or server can implement the method for generating music files provided in the embodiments of the present application by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can be a native application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a music APP or an instant messaging APP; it can also be a small program, that is, a program that can be run only by downloading it to a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be an application, module or plug-in in any form.

[0072] The method for generating a music file provided in the embodiment of the present application will be described in combination with the exemplary application and implementation of the terminal provided in the embodiment of the present application.

[0073] The following describes the method for generating a music file provided by the embodiment of the present application. As mentioned above, the electronic device for implementing the method for generating a music file in the embodiment of the present application can be a terminal or a server, or a combination of the two. Therefore, the execution subject of each step will not be repeatedly described below.

[0074] See also Figure 3A , Figure 3A is a flowchart of a method for generating a music file provided in an embodiment of the present application, which will be combined with Figure 3A The steps shown are explained.

[0075] In step 301, at least one modality data is acquired.

[0076] Here, each modality is associated with at least one musical element.

[0077] For example, modal data includes but is not limited to images, text, tags, videos, and audio (e.g., human voice humming). Each modal data can be converted to obtain the characteristics of at least one music element. Music elements refer to the various elements that constitute music. In the embodiment of the present application, emotion (P E )、Style(P G ), rhythm (P R ), Note(P N ) These four types of music elements are used as examples to illustrate.

[0078] In step 302, feature extraction processing is performed on at least one modal data to obtain at least one modal feature.

[0079] For example, when the modal data is an image, the image data can be feature extracted through a convolutional neural network to obtain the pixel value of each pixel in the image, thereby forming image features.

[0080] When the modal data is text, each character in the text can be encoded through the converter model to obtain text features. When the modal data is a label, the label can be combined into a feature sequence. When the modal data is a video, each video frame image in the video is extracted through a convolutional neural network to obtain a video feature composed of frame image features of multiple video frame images. When the modal data is audio, the audio data can be converted into audio features in the form of a sequence of notes.

[0081] In step 303, mapping processing is performed on each modal feature according to the type of the music element associated with each modal feature to obtain the element features of each music element associated with the modal feature.

[0082] For example, the corresponding relationship between the modal feature, the music element to be analyzed, and the feature of the music element can be represented as follows: the music element corresponding to the image modality is the emotion element {P E The music element corresponding to the text mode is the emotional element {P E}. The music element corresponding to the label mode is the emotional element {P E} and genre elements {P G The music element corresponding to the video mode is the emotional element {P E} and rhythmic elements {P E The musical element corresponding to the humming mode is the rhythmic element {P R} and note elements {P N The mapping processing methods for converting features of different modalities into corresponding elements are different, which are explained in detail below.

[0083] In some embodiments, the modality includes an image modality; the type of music element associated with the image modality is an emotional element; step 303 can be implemented in the following manner: performing emotional type prediction processing on the modal features of the image modality to obtain a first prediction probability that the modal features belong to different emotional types; and using the emotional type corresponding to the highest first prediction probability as the element feature of the emotional element.

[0084] For example, the emotion type prediction process can be implemented by a classification model, and the emotion types include but are not limited to: joy, laziness, sadness, playfulness, excitement, romance, tranquility, horror, and grandeur.

[0085] In some embodiments, the modality includes a text modality; the type of the music element associated with the text modality is an emotional element; step 303 may be implemented in the following manner: determining the element feature of the emotional element of the text modality in any of the following manners:

[0086] Method 1: Perform emotion type prediction processing on the modal features of the text modality to obtain the second prediction probability that the modal features belong to different emotion types, and use the emotion type corresponding to the highest second prediction probability as the element feature of the emotion element.

[0087] For example, the emotion type prediction process can be implemented through a classification model, which has been explained above and will not be repeated here.

[0088] Method 2: Obtain the first similarities corresponding to the text features of multiple reference emotion texts and the modal features of the text modality, and use the emotion type to which the reference emotion text corresponding to the highest first similarity belongs as the element feature of the emotion element, wherein each reference emotion text belongs to a different emotion type.

[0089] For example, the difference between the length of the reference emotion text and the length of the text used to generate the music file is less than the difference threshold, and the two are written in the same language. The method of extracting the text features of the reference emotion text is the same as the method of obtaining the modal features. The method of obtaining the first similarity can be to calculate the cosine similarity between the features.

[0090] In some embodiments, the modality includes a tag modality; the types of music elements associated with the tag modality include: emotional elements and style elements; step 303 can be implemented in the following manner: obtaining a mapping relationship table between tags and music elements, wherein the mapping relationship table includes: a mapping relationship between different tags and different types of emotional elements, and a mapping relationship between different tags and different types of style elements; according to the mapping relationship table, mapping the modal features of the tag modality to element features of at least one of the emotional elements and the style elements.

[0091] For example, the mapping relationship table is preset, and the mapping relationship can be determined through experience.

[0092] In some embodiments, the modality includes a video modality; the types of music elements associated with the video modality include: rhythmic elements and emotional elements; the modal features of the video modality include: video frame sequence, scene switching rate of the video frame sequence, and average optical flow intensity.

[0093] refer to Figure 3B , Figure 3B It is a flowchart of a method for generating a music file provided in an embodiment of the present application; Figure 3A Step 303 can be achieved by Figure 3B Steps 3031 to 3036 are implemented as described in detail below.

[0094] In step 3031, the music speed is determined based on the scene switching rate and the tangent function.

[0095] For example, the music tempo represents the number of beats per minute in the music file. scene is the total number of scene switches in the video N scene With video duration T video The ratio between them, scene switching rate R scene It can be expressed as the following formula (4.1):

[0096]

[0097] The music speed can be expressed as the following formula (4.2):

[0098] t music =t init +t inc *tanh(R scene ) (4.2)

[0099] Among them, the total number of scene switching in the video is N scene It can be calculated by the existing scene detection algorithm; video duration T video The unit is seconds; the scene switching rate R scene is the average number of scene changes per second in the video. Music speed t music Beats per minute (bpm) is used as the speed unit. init is the predetermined initial speed, t inc is the incremental speed, t music According to the scene switching rate R scene The music speed is calculated. The coefficient of the tanh activation function limits the incremental speed to between 0 and 1. The embodiment of the present application takes the predetermined initial speed t init =60, incremental speed tinc =70. The parser will get the music speed t music The calculation result is analyzed into the music speed (tempo) in the rhythmic element.

[0100] In step 3032, the ratio of the average optical flow intensity to the video size is used as the optical flow amplitude value of the video, and the ratio of the optical flow amplitude value to the number of music frames is used as the note density.

[0101] For example, the average light flux intensity is the light intensity per unit area of ​​the luminous surface, which can be expressed as the following formula (6.1):

[0102]

[0103] Among them, flow t (x, y) is the optical flow, H and W are the height and width of the video respectively. Number of music frames N fpb It can be determined by the following formula (6.2):

[0104]

[0105] Among them, fp svideo is the video frame rate, T video is the video duration, N bar is the number of measures of music.

[0106] The note density can be expressed as the following formula (6.3):

[0107]

[0108] Among them, density i N is the note density of a measure of music. fpb is the total number of video frames in a video clip. The duration of a video clip can be equal to the duration of at least one music measure.

[0109] In step 3033, the instantaneous optical flow change rate of the video frame sequence is used as the beat intensity.

[0110] For example, the instantaneous rate of change of light flow can represent the strength of the beat. i;j ) is calculated by the intra-beat visual beat saliency. Step 3033 can be represented by the following formula (6.4):

[0111]

[0112] Among them, strength i;j is the beat intensity, is the instantaneous rate of change of light current.

[0113] In step 3034, music speed, note density, and beat intensity are combined into element features of the rhythmic element.

[0114] For example, in the combination process, the order of music speed, note density and beat intensity is not important. The element features can be represented as feature sequences or matrices.

[0115] In step 3035, emotion type prediction processing is performed on the video frame sequence to obtain a third prediction probability that the modal feature belongs to a different emotion type.

[0116] For example, the emotion type prediction process can be implemented through a classification model, which will not be described in detail here.

[0117] In step 3036, the emotion type corresponding to the highest third prediction probability is used as the element feature of the emotion element.

[0118] For example, if the video is divided into multiple video clips, the emotion type prediction process for the video frame sequence is performed at the video clip level, and each video clip corresponds to at least one music section. The calculation formula includes the following formula (5.1), formula (5.2), and formula (5.3):

[0119]

[0120]

[0121]

[0122] Among them, the preset number of frames N is uniformly extracted in each section ipb frames (the number of images per bar), and calculate the emotion score S for each frame e (image), and average them to get the section sentiment score S e (bar). bpb is the number of beats per measure, combined with the video length T video and music speed music The number of music measures N can be calculated bar . All N bar Section emotional score S e (bar) is averaged to get the video emotion score S e In the embodiment of the present application, a preset number of frames N is uniformly extracted from each section. ipb Equal to 8, the number of beats per measure is N bpb Take 4.

[0123] In some embodiments, the types of emotions that can be extracted from the video are the same as other modalities. If the video itself carries audio, the audio of the video can be separated from the image, and the audio can be used as data of the audio modality for music generation.

[0124] In some embodiments, the modality includes an audio modality; the types of music elements associated with the audio modality include: rhythm elements and note elements; Figure 3C , Figure 3C It is a flowchart of a method for generating a music file provided in an embodiment of the present application; Figure 3A Step 303 can be achieved by Figure 3C Steps 3037 to 3038 are implemented as described in detail below.

[0125] In step 3037, the modal features of the audio modality are transcribed to obtain the element features of the note elements.

[0126] For example, the element features of the note element are represented by a first note sequence, and the first note sequence includes multiple notes and the playback order corresponding to each note. The transcription process is to parse the modal features of the audio mode into a sequence composed of notes, and the file obtained by transcribing the audio file is in MIDI format. The data of the audio mode can be obtained by recording human humming. The transcription process can be implemented by the VOCANO algorithm. The transcribed audio file can also be standardized to improve the accuracy of the generated music.

[0127] For example, using the note transcription framework of singing voice in polyphonic music using the VOCANO algorithm, the input humming audio file is parsed to obtain the humming melody, and a five-line score or a simplified score is generated to transcribe the original preamble M origin , and then use the standardization module to convert the original preamble M origin Processed into standard preamble M std This process is expressed as the following formula (8.1) and formula (8.2):

[0128] M origin =VOCANO(humming) (8.1)

[0129] M std =Standardize(M origin ) (8.2)

[0130] Among them, humming is the modal feature corresponding to humming.

[0131] In step 3038, feature extraction processing is performed on the first note sequence to obtain element features of the rhythmic elements.

[0132] For example, the element features of the rhythm element include music speed, note density and beat intensity. std Analyze the corresponding note elements and rhythm elements in the music element projection space. The formula is expressed as the following formula (9.1) to formula (9.3):

[0133]

[0134]

[0135]

[0136] Where i = 1, 2, ..., N bar and j=1, 2, ..., N bpb ; F XP represents the parser, humming is the input humming audio content. D and S represent the calculation formulas of note density and beat strength respectively. Represents the standard prefix M std The fixed beat length in seconds. Represents the standard prefix M std The i-th measure, the j-th beat of the i-th measure, and the tempo value of the beat are used to calculate the rhythmic elements.

[0137] In the embodiments of the present application, it supports converting data of different modalities into music elements, which improves the freedom of generating music files. Compared with the solution of directly converting modal data into audio data, using music elements as intermediate values ​​can improve the accuracy of converted music files and improve the degree of matching between music files and original input data.

[0138] Continue to refer Figure 3A , in step 304, each element feature is mapped into a music representation sequence.

[0139] For example, the music representation sequence includes the sound information corresponding to different moments in the music file. Figure 7 , Figure 7 Detailed description of the structure of the music representation sequence provided in the embodiment of the present application. On the time axis of the music representation sequence, specific parameters corresponding to different music elements at each moment are arranged in sequence.

[0140] In some embodiments, the types of music elements include: note elements, style elements, emotional elements, and rhythmic elements; Figure 3D , Figure 3D is a flowchart of a method for generating a music file provided in an embodiment of the present application; Figure 3AStep 304 can be accomplished by Figure 3D Steps 3041 to 3043 are implemented as described in detail below.

[0141] In step 3041, a second note sequence of at least one audio track is generated based on the element features of the note elements.

[0142] For example, the second note sequence includes the order in which each note is played. In audio editing software, the audio can be composed of sub-files of different tracks (corresponding to the second note sequence), and the sub-files of multiple tracks are superimposed to form a final music file. The types of tracks can include accompaniment tracks and melody tracks.

[0143] In step 3042, parameters of each second note sequence are configured based on the rhythmic elements to obtain a third note sequence.

[0144] For example, the rhythmic elements include the note density and beat strength of each measure, and the time point and strength corresponding to each note of the second note sequence in the time axis are configured based on the rhythmic elements to form a third note sequence.

[0145] In step 3043, each third note sequence is labeled based on the element features of the emotional elements and the musical style elements, and the labeled third note sequences are superimposed to obtain a music representation sequence.

[0146] Here, the music representation sequence includes the sound information of each playing moment, and the sound information includes: music style type, emotion type, notes, and the note density and beat intensity corresponding to the notes. The third note sequence is superimposed to form Figure 7 The structure shown is a sequence of music representations in matrix form.

[0147] In the embodiment of the present application, by generating note sequences corresponding to different tracks and forming a final music representation sequence, the music file converted based on the music representation sequence can be made more vivid and three-dimensional, thereby improving the quality of the generated music file.

[0148] Continue to refer Figure 3A In step 305, the music representation sequence is decoded to obtain a music file.

[0149] For example, the decoding process includes continuing to write the subsequent sound information in the music representation sequence based on the music representation sequence, and decoding the continued music representation sequence from the feature form into a music file in MIDI format. The music file in MIDI format can be played by calling a sound card.

[0150] In some embodiments, the music representation sequence includes the sound information at each moment; the sound information is represented by a word matrix composed of different information; Figure 3E , Figure 3Eis a flowchart of a method for generating a music file provided in an embodiment of the present application; Figure 3A Step 305 can be accomplished by Figure 3E Steps 3051 to 3053 are implemented as described in detail below.

[0151] In step 3051, the word unit matrix of the sound information of the music representation sequence is aggregated to obtain a word unit encoding vector.

[0152] For example, the aggregation process can be implemented through linear transformation and position encoding processing to aggregate the information in the word unit matrix and reduce the dimension.

[0153] In some embodiments, step 3051 can be implemented in the following manner: perform linear mapping processing on the word unit matrix of each sound information in the music representation sequence to obtain a coding embedding vector; connect each coding embedding vector according to time sequence to obtain a concatenated feature vector; perform linear mapping on the concatenated feature vector, and perform position encoding processing on the result of the linear mapping to obtain a word unit coding vector.

[0154] For example, the sound information is represented by a token matrix composed of different information, each of which can be regarded as a token. The token matrix representation is a two-dimensional event matrix, whose matrix elements Represents the jth attribute value of the i-th word. Here eventi has a total of 12 dimensions, which represent the values ​​of the current word in the attributes of "family word type", "emotion", "music style", "bar / beat", "speed", "chord", "density", "intensity", "number", "pitch", "duration", "intensity" and so on. For time i, each matrix element are linearly mapped to the encoding embedding vector It can be represented by the following formula (10.1):

[0155]

[0156] Where i is any value from 1 to t. According to the example above, j is any value from 1 to 12. Onehot is a one-hot vector conversion function. Linear is a linear conversion function.

[0157] Embed all the codes into vector Connect them in order according to the generation time to get the concatenated feature vector of the current family word i , concatenate feature vectors i Perform linear mapping and position encoding in sequence to obtain the word unit encoding vector Input i The above process can be represented by the following formula (10.2) to formula (10.3):

[0158]

[0159] Input i =Linear input (Positionalencoding(concat i )) (10.3)

[0160] Among them, Positionalencoding is the position encoding function.

[0161] In step 3052, multiple rounds of sound information prediction processing are performed based on the word-unit encoding vector of the music representation sequence to obtain an updated music representation sequence.

[0162] For example, the updated music representation sequence is composed of the music representation sequence and a plurality of predicted sound information. The number of rounds can be preset, and the sound information prediction processing is used to predict the sound information of the current round based on the predicted sound information and the music representation sequence predicted in the previous round.

[0163] In some embodiments, step 3052 can be implemented in the following manner: taking the music representation sequence as the first music representation sequence; predicting the predicted sound information of the next moment corresponding to the first music representation sequence based on the word unit encoding vector of the current music representation sequence; adding the predicted sound information to the end of the current music representation sequence to obtain a second music representation sequence; taking the second music representation sequence as the first music representation sequence, and performing processing to predict the predicted sound information of the next moment corresponding to the first music representation sequence based on the word unit encoding vector.

[0164] For example, the prediction process can be implemented by a transformer model, and the music representation sequence is processed in the word dimension. The above steps can be represented as the following formulas (10.4) to (10.8):

[0165]

[0166]

[0167]

[0168]

[0169]

[0170] Among them, H t is the latent variable at the current time t obtained by the transformer model decoder (Transformerdecoder), FT t is the hidden variable Ht The linear transformation result. Yes FT t , H t The result obtained by linear transformation of the splicing result. Through the FT t or Obtained by calling the function argmax with the independent variable.

[0171] In some embodiments, in response to the length of the second music representation sequence reaching a preconfigured length, the second music representation sequence is used as an updated music representation sequence. If the length of the second music representation sequence is equal to the preconfigured length, the prediction process is stopped and the current second music representation sequence is used as an updated music representation sequence.

[0172] In step 3053, the updated music representation sequence is format-converted to obtain a music file.

[0173] For example, the format of the music file is MIDI format, which can be used to call the sound card to play the sound information.

[0174] In the embodiment of the present application, multiple rounds of sound information prediction are performed based on the existing music representation sequence, which improves the correlation between the contexts of the sound information in the music file, thereby obtaining more logical and regular music, and improving the accuracy of generated music.

[0175] In some embodiments, the number of music files is multiple; Figure 3F , Figure 3F is a flowchart of a method for generating a music file provided in an embodiment of the present application; after step 305, execute Figure 3F Steps 306 to 308 are described in detail below.

[0176] In step 306, the following processing is performed on each music file: the sound information of each moment of the music representation sequence is predicted to obtain the prediction probability corresponding to the music representation sequence in different evaluation dimensions.

[0177] For example, the evaluation dimensions include: music quality, music emotion, and music style. Taking the music quality dimension as an example, the prediction probability includes the probability that the music representation sequence belongs to high quality and low quality types respectively. Taking the music emotion type as an example, the prediction probability includes the probability that the music representation sequence belongs to different types of emotions; taking the music style type as an example, the prediction probability includes the probability that the music representation sequence belongs to different types of styles. The prediction process can be implemented by a classification model.

[0178] In some embodiments, the sound information is represented by a word-gram matrix; step 306 can be implemented in the following ways: perform encoding processing on the word-gram matrix of each sound information to obtain the encoding features of each sound information; perform global average pooling processing on each encoding feature according to the time dimension to obtain the global feature; normalize the global feature for each evaluation dimension to obtain the predicted probability of the music representation sequence corresponding to each evaluation dimension.

[0179] In step 307, a weighted sum is performed on each prediction probability to obtain a quality index of the music file.

[0180] For example, the weight values ​​used in the weighted summation process are preset according to the degree of attention paid to different evaluation dimensions in the application scenario, for example, the weight value of the quality dimension is higher than that of other dimensions.

[0181] In some embodiments, other dimensions may be used as references, and the predicted probability of the quality dimension may be used as an instruction indicator.

[0182] In step 308, the music file corresponding to the highest quality index is used as the optimal music file.

[0183] For example, each quality index is sorted in descending order to obtain the highest quality index, and the music file corresponding to the highest quality index is used as the optimal music file, the optimal music file is retained, and other music files are deleted.

[0184] The embodiment of the present application reduces the number of low-quality files in the generated music files through quality evaluation, and performs quality evaluation through artificial intelligence without the need for human intervention, thereby saving the cost and time of evaluating music quality.

[0185] In some embodiments, reference Figure 4 , Figure 4 4 is a structural diagram of a music file model provided in an embodiment of the present application. The music file generation method is implemented by a music generation model 401, and the music generation model 401 includes an encoder 402 and a decoder 403; wherein the encoder 402 is used to obtain the following features: modal features and element features of each music element associated with the modal features, that is, to execute steps 302 to 303; the decoder 403 is used to obtain a music representation sequence, and to decode the music representation sequence, that is, to execute steps 304 to 305. The music generation model also includes a filter 404, and when there are multiple music files generated, the filter is used to determine the optimal music file from the multiple music files.

[0186] In the embodiment of the present application, the modal data is mapped to the element features corresponding to the music elements, the element features are mapped to the music representation sequence, and the music file is generated based on the music sequence. Multimodal controllable, supports the generation of music files through different modal data, and improves the freedom of obtaining music files compared to the solution of generating music files based on single modal data in the related technology; modal data is mapped to music elements, and then the music files are determined by the music elements, which improves the granularity of the music files generated by modal data. Compared with the solution of related technology relying on music materials, it saves computing resources and improves the matching degree between the input modal data and the output music files, so that the generated music files can be more accurate.

[0187] Next, an exemplary application of the method for generating a music file according to an embodiment of the present application in a practical application scenario will be described.

[0188] In the related art, music generation schemes based on artificial intelligence can be roughly divided into five categories: uncontrolled, label-controllable, sequence-controllable, video-controllable and multimodal-controllable.

[0189] (1) Uncontrolled music generation. Uncontrolled music generation schemes use random seeds to generate music from scratch without any specific constraints or additional inputs. The most critical challenge is how to ensure long-term structural consistency as the length of the music increases. In the prior art, the overall repetitive structure of the generated music is enhanced through transformer models, recurrent neural network (RNN) models, or optimization-based models. In contrast, controlled music generation has become increasingly popular in recent years because it allows users to provide input information to generate unique musical works.

[0190] (2) Label-controllable music generation. Label-controllable music generation schemes generate music based on high-level semantic labels (such as instruments, music styles, or emotions) as control conditions. For example, the online AI music generation tool MuseNet can generate conditional music based on a specific set of instruments and a specific music style.

[0191] (3) Sequence-controlled music generation. Existing technologies usually use a control sequence as a priori and generate a continuation of the sequence accordingly. The types of control sequences for music generation usually include scores, melodies, themes, motifs, and lyrics. These sequences can be directly extracted from music works and form training pairs with the original music.

[0192] (4) Video-controlled music generation. Existing video-controlled music generation schemes aim to create music from silent performance videos, which can also be regarded as visual-to-music transcription. The instrument type and music rhythm can be inferred from visual clues (e.g., performance venue, musician movements, etc.), which limits the diversity of music to some extent.

[0193] However, the prior art has the following problems

[0194] (1) The problem of a single controllable modality. Existing methods usually use labels, melodies, lyrics, videos, etc. as control conditions to generate music that matches the input modality. The control modalities of these methods are relatively simple, and due to the inherent differences between modalities, music generation solutions designed for a single modality are difficult to directly migrate to other modalities. For example, music generation models trained for videos are difficult to directly migrate to text, humming, etc. (2) The problem of low matching between generated music and the control modality. Conventional methods calculate similarity based on modal features and music clip features, and use the most similar music clip as the soundtrack result of the modality. This is heavily dependent on the richness of music clips in the music library, and the generated music has a low matching degree with the input modality in terms of rhythm and emotion (rough matching at the input modality and audio clip level). (3) The problem of unstable quality of generated music due to the lack of quality evaluation. Mainstream music generation solutions directly use the output of the AI ​​model as the final result, which inevitably results in some poor quality content that is lower than the average level of the model. (4) The problem of poor scalability of the model control modality. Mainstream methods are designed for specific modalities and are difficult to directly expand to new modalities.

[0195] In summary, the existing technical solutions have the problems of single controllable mode, low matching degree between generated music and control mode, unstable quality of generated music due to lack of quality assessment, and poor scalability of model control mode. In view of the above problems, the embodiment of the present application proposes a music generation solution based on multimodal analysis and control, which can analyze multimodal input data to control the music generation model to generate music that matches the input modality emotion, style, and rhythm. Specifically, by understanding the content of multimodal data (pictures, videos, labels, text, humming), and parsing the understanding results into the music element projection space (emotion, style, rhythm, notes, etc.), and based on the projection results (audio elements), the music generation model is controlled to generate an audio feature sequence that matches it, and the audio feature sequence is restored to a music file, so as to solve the problems of single controllable mode, low matching degree between generated music and control mode, and poor scalability in the existing music generation method.

[0196] In the embodiment of the present application, according to the modal content analysis, parsing to the music elements, mapping to the music representation sequence, controlling the generation of matching music, and controlling the generation strategy in a fine-grained manner, the generated music and the input modality are highly matched in terms of the four types of music elements: emotion, style, rhythm, and notes. By introducing a supervised multi-task filter to control the quality of the output music, the design of the multi-task can guide the network to pay attention to the emotion and style information while judging the quality. Only the samples that meet the quality score passing line and have the highest score will be used as the output result, thereby achieving high-quality music generation. The embodiment of the present application decouples the parsing of the multi-modal input control signal from the music control generation part, that is, the training and control of the music generation model do not depend on a specific modality. The advantage of this design is that it is easy to "plug-in" the newly added modality. For example, if you want to add a new modality of "human posture", analyze the characteristics of the newly added modality through the parser, and parse it to a specific music element, and use the music element as a control signal to control it, without retraining the music generator, and it has strong scalability and generalization.

[0197] With the server as the execution body, combined with Figure 5 The steps of the present invention are used to illustrate the method for generating a music file provided in the embodiment of the present application. Figure 5 , Figure 5 It is an optional flow chart of the method for generating a music file provided in an embodiment of the present application.

[0198] In step 501, multimodal data is acquired.

[0199] For the convenience of explanation, the principle of the method for generating a music file provided in the embodiment of the present application is explained. Figure 8 , Figure 8 It is a schematic diagram of the principle of the method for generating a music file provided in an embodiment of the present application.

[0200] Parser 601 converts multimodal data into music elements, generator 602 converts music elements into characterization sequences, decoder 6021 converts characterization sequences into music files, and filter 603 filters the best music files. The present application embodiment supports the content of various modes as input prompts (prompt), and can control the generation of high-quality music that matches it. The present application embodiment is divided into three links: parsing, generating, and screening. First, the parser will analyze the input modal content and parse it into the music element projection space. Then, the music elements will be mapped to the music characterization sequence, and then the music generator will be controlled to generate matching music. Finally, the filter evaluates the quality of the generated music and filters out the music with the highest quality.

[0201] In step 502, the multimodal data is parsed into features corresponding to the music elements.

[0202] For example, the types of multimodal data include images, text, tags, videos, and human voice humming, etc. Each modal data can be converted into the characteristics of at least one music element. The symbolic music element projection space, defined as P, is a bridge connecting multimodal content and symbolic music. It contains emotions (P E )、Style(P G ), rhythm (P R ), Note(P N ) These four musical elements are: E , P G , P R , P N}∈P.

[0203] Emotion and style elements are elements covering the entire range of music, and the features of emotion elements and style elements are all one-hot vectors. In the embodiment of the present application, emotions include but are not limited to 9 categories: joy, laziness, sadness, playfulness, excitement, romance, quietness, horror, and grandeur, and styles include but are not limited to 5 categories: pop, classical, Chinese style, jazz, and light music.

[0204] Rhythmic elements (P R ), whose components cover the subsection range and can be expressed as Here N bar is the number of measures in the whole song. represents the rhythmic component of the i-th measure, which can be expanded into Here I n Indicates the number of beats in the ith measure. More specifically, measure element p bar =(bar,density)∈R 2 Record the starting position (bar) and note density (density) of the current measure, beat element p beat =(beta,tempo,strength)∈R 3 Record the start position (beat), speed (tempo) and beat strength (strength) of the current beat.

[0205] Note element (P N ), whose components cover the range of a single note and can be expressed as Here N note is the number of notes in a note sequence. n =(pitch,duration,velocity)∈R 3 Indicates the pitch, duration, velocity, and other information of the current note.

[0206] When prompt words of different modalities are input, the modal content will be analyzed and parsed into specific musical elements. These parsed musical elements will serve as control conditions to guide the generation of music matching the modality. The corresponding relationship between the input modality, the musical elements to be analyzed, and the characteristics of the musical elements can be represented as follows:

[0207] The music element corresponding to the image is the emotional element {P E}.

[0208] The music element corresponding to the text is the emotional element {P E}.

[0209] The music element corresponding to the label is the emotional element {P E} and genre elements {P G}.

[0210] The music element corresponding to the video is the emotional element {P E} and rhythmic elements {P R}.

[0211] The musical element corresponding to humming is the rhythmic element {P R} and note elements {P N}.

[0212] For example, the above five modalities are used as examples in the embodiments of the present application. Other modalities (including but not limited to: line images, human body postures, depth images, infrared images, etc.) should also be included in the protection scope of the embodiments of the present application as long as they adopt a similar solution of "parsing music elements as the basis for generating music files".

[0213] For the sake of explanation, the structure of the parser proposed above is explained, refer to Figure 6 , Figure 6 601 includes: a picture emotion analysis module 6011, a text emotion analysis module 6012, a label mapping module 6013, a video content analysis module 6014, a transcription module 6015, and a standardization module 6016. The picture emotion analysis module 6011 is used to convert pictures into emotion elements; the text emotion analysis module 6012 is used to convert text into emotion elements; the label mapping module 6013 converts labels into music style or inclination elements; the video content analysis module 6014 is used to convert the scene switching rate and motion information of the video into rhythmic elements, and convert the bar and video-level emotions into emotion elements; the transcription module 6015 and the standardization module 6016 are used to convert the humming audio into note elements and rhythmic elements.

[0214] For the picture modality, the embodiment of the present application uses the picture emotion analysis module to identify the emotion category of the picture, that is, to calculate the emotion score S of the picture.e (image), take the emotion with the highest score as the analysis result, and parse it into the specified emotion element in the music element projection space, so as to control the music generation of the specified emotion. It is expressed as the following formula (1):

[0215] F XP (image) = P E ={arg max e∈ε S e (image)} (1)

[0216] Among them, F XP represents the parser, and image is the input image. arg stands for argument, which means "independent variable" here. argmax is a function that finds the parameter (set) of a function. In argmaxg(t), it expresses a subset of the domain, and any element in the subset can make the function g(t) take the maximum value. The sentiment analysis module can be a classification model, such as: convolutional neural network, residual network (ResNet), VGG, etc.

[0217] For the text modality, the embodiment of the present application identifies the emotion category of the text through the text emotion analysis module, and parses it into the same emotion element in the music element projection space, thereby controlling the music generation. The above relationship can be expressed as the following formula (2):

[0218]

[0219] Among them, F XP represents the parser, and text is the input text. The sentiment score S of the text can be directly identified by the text classification model e (text), take the emotion with the highest score as the text emotion analysis result; you can also use a text encoding model (such as Sentence converter, etc.) to calculate the feature similarity score S of the input text and the text sentence with synonyms of emotion e e (text), take the emotion with the highest similarity score as the text sentiment analysis result.

[0220] For the tag mode, the embodiment of the present application supports inputting emotion or genre tags to control music generation. When the user inputs an emotion (such as joy) or a genre tag (such as pop), the parser directly parses the tag to the same emotion or genre element in the music element projection space through the tag mapping module. The above mapping relationship can be expressed as the following formula (3.1) and formula (3.2):

[0221] F XP (tag e )=P E ={tag e} (3.1)

[0222] F XP (tag g )=P G ={tag g} (3.2)

[0223] Among them, F XP Represents a parser, tag e is the input emotion label, tag g The input genre label.

[0224] For video modality, the embodiment of the present application analyzes the scene switching rate, emotion, motion and other information of the video through the video content analysis module, and parses this information into corresponding rhythmic elements and emotional elements in the music element projection space through the parser to generate music that highly matches the video. The video content analysis module can be a converter model or a model for video feature extraction.

[0225] The scene switching rate is analyzed to the tempo in the rhythmic element. Based on experimental observations, the music speed of high-quality video background music is usually related to the video scene switching frequency. For example, when multiple clips are switched in a short period of time in a mixed-cut video, it is usually accompanied by fast music, while the scene switching frequency in a landscape video is low and is usually accompanied by gentle music. For this reason, the embodiment of the present application defines the video scene switching rate R scene Control music speed music , the calculation formula is represented by the following formulas (4.1) and (4.2):

[0226]

[0227] t music =t init +t inc *tanh(R seene ) (4.2)

[0228] Among them, N scene represents the total number of scene switches in the video, which can be calculated by the existing scene detection algorithm; T video Represents the video duration in seconds; R scene The average number of scene changes per second in the video. Music speed t music Beats per minute (bpm) is used as the speed unit. init is the predetermined initial speed, t inc is the incremental speed, t music According to R scene The music speed is calculated. The coefficient of the tanh activation function limits the incremental speed to between 0 and 1. The embodiment of the present application takes the predetermined initial speed tinit =60, incremental speed t inc =70. The parser will get the music speed t music The calculation result is analyzed into the tempo in the rhythmic element.

[0229] The results of the bar-level and video-level emotion recognition are parsed into emotion elements. The embodiment of the present application uses the video emotion analysis module to identify the emotion category of the input video and parse it into the same emotion element in the music element projection space, thereby controlling the music emotion. Similar to the picture emotion analysis module, the video emotion analysis module calculates the emotion score S of the video. e (video), the emotion with the highest score is taken as the analysis result, and its calculation formula includes the following formula (5.1), formula (5.2), and formula (5.3):

[0230]

[0231]

[0232]

[0233] Among them, the preset number of frames N is uniformly extracted in each section ipb frames (the number of images per bar), and calculate the emotion score S for each frame e (image), and average them to get the section sentiment score S e (bar). bpb is the number of beats per measure, combined with the video length T video and music speed music The number of music measures N can be calculated bar . All N bar Section emotional score S e (bar) is averaged to get the video emotion score S e In the embodiment of the present application, a preset number of frames N is uniformly extracted from each section. ipb Equal to 8, the number of beats per measure is N bpb Take 4.

[0234] The motion information is parsed into the note density and beat strength in the rhythmic elements. Fast motion frames should correspond to dense notes. The embodiment of the present application uses video motion information to control the local rhythm of the music. Specifically, the video motion analysis module calculates the optical flow t(x, y) to obtain video motion information, and then the parser parses the average optical flow intensity in the measure (the average amplitude of the optical flow in the measure) and the visual beat significance in the beat (the instantaneous optical flow change rate) into the bar note density (density) and beat strength (strength) of the same percentile in the music element projection space. The specific calculation formulas include the following formulas (6.1) to (6.4):

[0235]

[0236]

[0237]

[0238]

[0239] Among them, H and W are the height and width of the video, F t is the optical flow amplitude value of frame t; N fpb The number of frames per section, fps video is the frame rate of the video, N bar is the number of musical measures; the density of notes i ) is obtained by the average optical flow intensity in the measure (the average amplitude of the optical flow in the measure), and the beat strength (strength i;j ) is obtained from the intra-beat visual beat salience (instantaneous rate of change of light flux). It is the calculation method of instantaneous rate of change of light flux.

[0240] Optical flow calculation can be achieved in the following ways: Brightness gradient-based optical flow calculation method uses the brightness gradient of pixels in the image to infer the movement of the object. By calculating the gradient vector of the pixel, the speed and direction of the object can be obtained.

[0241] At this point, the above video content analysis module has completed the analysis of the scene switching rate, emotion, movement and other information of the video. The parser then parses this information into the corresponding rhythmic elements and emotional elements in the music element projection space. The complete mapping relationship can be expressed as the following formulas (7.1) to (7.3):

[0242]

[0243]

[0244]

[0245] Among them, F XP represents the parser, video is the input video. i=1,2,...,N bar and j=1, 2, ..., Nbpb .

[0246] For the humming mode, the embodiment of the present application guides the generation of complete music by transcribing the humming audio data into pre-order MIDI. Specifically, a transcription module (such as the note transcription framework of singing sound in VOCANO algorithm polyphonic music) is used to parse the input humming audio file to obtain the humming melody, and generate a five-line score or a simplified score to transcribe it into the original pre-order MIDI . origin , and then use the standardization module (Standardize) to convert the original preamble M origin Processed into standard preamble MIDIM std This process is expressed as the following formula (8.1) and formula (8.2):

[0247] M origin =VOCANO(humming) (8.1)

[0248] M std =Standardize(M origin ) (8.2)

[0249] The parser adds the standard prefix M std Analyze the corresponding note elements and rhythm elements in the music element projection space to continue the music sequence. The formula is expressed as the following formula (9.1) to formula (9.3):

[0250]

[0251]

[0252]

[0253] Where i = 1, 2, ..., N bar and j=1, 2, ..., N bpb ; F XP represents the parser, humming is the input humming audio content. D and S represent the calculation formulas of note density and beat strength respectively. Represents the standard prefix M std The fixed beat length in seconds. Represents the standard prefix M std The i-th measure, the j-th beat of the i-th measure, and the tempo value of the beat are used to calculate the rhythmic elements.

[0254] In step 503, the features corresponding to the music elements are mapped into a music representation sequence.

[0255] refer to Figure 7, Figure 7 Schematic diagram of the structure of the music representation sequence provided by the embodiment of the present application. The original symbolic music or the music elements parsed by the parser are mapped to the symbolic music representation sequence, thereby assisting the training, reasoning control, quality assessment and screening process of the music generation model. The music elements will be mapped to the music representation sequence, and the music generator can be called to generate matching music based on the music representation sequence.

[0256] There is no limitation on the order of family units and units in the music representation sequence. For example, the "rhythm" family units include "beat", "speed", "chord", "intensity" and other units, and the order of these units is not limited. The "number" unit in the "track" family unit is not limited to the two tracks of melody and accompaniment. For multi-instrument music, the track numbers here can be different instruments. The decoder is not limited to the converter decoder, and any model that can perform sequence prediction can be used to implement the embodiments of the present application.

[0257] For example, in order to obtain a music representation sequence, the embodiment of the present application adopts the architecture of a compound words converter, that is, tokens belonging to the same family are grouped into a super token and placed at the same time position. Tokens of the same family belong to the same event type.

[0258] The Compound Words converter is an improvement based on the converter model, and is used to realize the composition of complete music based on dynamic directed hypergraphs. A hypergraph is a generalized graph, characterized by a hyperedge that can connect multiple points. A hypergraph H is a set H = (X, E), where X is the set of vertices and E is the non-empty power set of X.

[0259] Based on the combination word converter, the embodiment of the present application introduces a new family token called "label" and two corresponding tokens ("emotion" and "style") to control the generation of music with specified emotions and styles. A new family token called "track" and the corresponding token "track number" (referred to as "number") are introduced to distinguish different instruments or accompaniment and melody tracks, thereby ensuring that the generated music has beautiful melody and smooth accompaniment. In the rhythm family token, tokens such as "note density" (referred to as "density") and "beat intensity" (referred to as "intensity") are added to the "bar" and "beat" events respectively to control the generation of music with specified beat density and note intensity.

[0260] Here, the embodiment of the present application uses the "label" family word-element to express the semantic information of the entire music. Fine-grained "label" family word-element is introduced, that is, the label family word-element is specified at the beginning of each measure. The advantage of this is that music that is more in line with the emotion and style labels can be generated, and in some specific task scenarios (such as: generating music from videos), it supports fine-tuning the emotion category measure by measure to ensure the generation effect. The "emotion" and "style" word-element represent the emotion and style information of the music respectively, and there are 9 and 5 options respectively.

[0261] The "track" family word element is located at the intersection of different tracks in a series of notes and is used to express the local track information of the music. In the data processing stage, the embodiment of the present application divides the notes in the MIDI file into two tracks according to melody and accompaniment. Therefore, the "number" word element has two options: melody and accompaniment, indicating whether the next series of notes belongs to the melody track or the accompaniment track. Through such a representation method, the embodiment of the present application realizes track-level modeling of melody and accompaniment to generate music with beautiful melody and smooth accompaniment. This is not limited to the two tracks of melody and accompaniment. For multi-instrument music, the track numbers here can be different instruments.

[0262] In the scenario of video-generated music, the rhythm of the generated music is controlled by the video motion information. To this end, the embodiment of the present application adds "density" and "intensity" words to the "bar" and "beat" word positions, respectively, to control the note density and beat intensity of the current bar, and generate music with a rhythm that matches the video.

[0263] The representation method of the embodiment of the present application records the symbolic music elements or the words such as "label", "bar", "beat", "track", "note" contained in each "bar" in the MIDI file in chronological order to form a representation sequence to assist the subsequent music generation model training, reasoning control, quality evaluation and screening process.

[0264] In step 504, a plurality of candidate music files are generated based on the music representation sequence.

[0265] The decoder is used to generate multimodal and controllable symbolic music. The decoder uses the decoder of the transformer model as the backbone network to learn the dependencies between word units.

[0266] Specifically, assuming that the first t tokens are known and we want to predict the value of the next token, we can do this in the following way:

[0267] Transform the word sequence into an equivalent two-dimensional event matrix, whose matrix elements are Represents the jth attribute value of the i-th word. Here event iThere are 12 dimensions in total, representing the values ​​of the current word in terms of "family word type", "emotion", "music style", "bar / beat", "speed", "chord", "density", "intensity", "number", "pitch", "duration", "intensity" and other attributes.

[0268] For time i, each matrix element are linearly mapped to dense vectors (the encoding embedding vector above), and concatenate the dense vectors of all elements in sequence to obtain the dense representation of the current family word i (the concatenated feature vector above), for dense representation concat i Performing linear mapping and position encoding in sequence will obtain the input feature Input of the converter network at time i i (The word encoding vector above).

[0269] The above process can be represented by the following formula (10.1) to formula (10.3):

[0270]

[0271]

[0272] Input i =Linear input (Positionalencoding(concat i )) (10.3)

[0273] Wherein, i is any value from 1 to t, and j is any value from 1 to 12.

[0274] Input the input features of the previous t moments into the converter network and obtain the hidden variable H at the current moment t , and by H t Perform multiple linear mappings to predict the next event i+1 Here, we first predict the "family word type" at the next moment, and then predict other words at the next moment based on this. The above process can be represented by the following formulas (10.4) to (10.8):

[0275]

[0276]

[0277]

[0278]

[0279]

[0280] Finally, the predicted two-dimensional event matrix is ​​reversed into MIDI according to the music representation to obtain the candidate music generation result.

[0281] In step 505, a target music file is screened from a plurality of candidate music files.

[0282] For example, multi-task learning in the filter is not limited to the three subtasks of quality assessment, emotion recognition, and genre recognition, but the quality assessment task is necessary. The generator uses a decoder to learn the dependencies between different word units based on the obtained symbolic music representation sequence during training. During reasoning, the parser uses the music elements mapped after multimodal analysis as control signals to predict the word units at the next moment, and reverse-translates the predicted two-dimensional event matrix into MIDI according to the music representation, thereby obtaining the candidate music generation results. The embodiment of the present application proposes that the filter perform quality assessment on the candidate music generation results and select the music with the highest quality as the final output result.

[0283] The filter identifies high-quality candidate music by constructing a multi-task learning scheme. The filter uses a converter encoder as the backbone network to determine the quality of music. In practical applications, under given control conditions, only a part of the music generated by the generator meets the high-quality requirements. This part of the generated results is beautiful and coherent, with obvious melody fluctuations and alternating strong and weak rhythms. The embodiment of the present application hopes to accurately identify this part of high-quality music through supervised learning. The embodiment of the present application first uses a generator to generate a batch of music under various emotional and style control conditions, and organizes manual annotation of each piece of music to see whether it can meet the high-quality requirements. Finally, the filter trains a classification model based on the above-mentioned supervised data to judge the quality of any input music.

[0284] Specifically, the embodiment of the present application constructs a multi-task learning solution including quality assessment, emotion recognition, and music style recognition to complete music screening. For a midi file, the filter expresses it as a word sequence based on the representation method proposed in the embodiment of the present application, and then further translates it into an equivalent two-dimensional event matrix. The embodiment of the present application calculates the event vector event at each moment of the event matrix. i Perform the same operation as the generated model to obtain the input feature Input of the current converter network. i Since the filter analyzes the entire midi file, rather than predicting the next event like a generator, the present embodiment uses all the input features of T moments as Input` i At the same time, it is sent to the encoder model to obtain the output features F at each moment i Next, the overall output feature F iPerform global average pooling along the time dimension to obtain the global feature F encoder Finally, the global feature F encoder By connecting to three fully connected layers and normalizing them, we can get the category probabilities of each task. The above process can be represented by the following formulas (11.1) to (11.5):

[0285]

[0286]

[0287] P genre =Softmax(FC genre (F encoder )) (11.3)

[0288] P emotion =Softmax(FC emotion (F encoder )) (11.4)

[0289] P quality =Softmax(FC quality (F encoder )) (11.5)

[0290] The reason for constructing a multi-task learning scheme here is that different types of music have slightly different standards for quality. The multi-task design can guide the network to pay attention to emotions while judging quality. emotion and genre information genre , improving the overall screening effect. In addition, although the training data comes from the music generated by the generator, the filter also has good generalization for symbolic music from other unknown sources and can accurately judge its quality, emotion and style information.

[0291] In the inference phase, the filter converts P quality The probability of medium and high quality categories is used as the quality score, and the samples with quality better than the passing line and the highest score in a batch are selected for output, thereby achieving high-quality music screening.

[0292] In some embodiments, the method for generating music files provided by the embodiments of the present application can be applied to multiple projects and product applications including video soundtracks, interactive entertainment, auxiliary creation, music education, music therapy, and personalized music generation, and can generate high-quality music with beautiful melody and matching emotions based on multimodal information input by users, thereby improving user experience. In the application scenario of video soundtracks, the embodiments of the present application have strong flexibility and scalability, can support multimodal (picture, video, label, text, humming) input, and can generate high-quality music that matches it under different modal controls.

[0293] The beneficial effects of the method for generating a music file provided in the embodiment of the present application include:

[0294] (1) Multi-modal controllability. Related technologies only focus on single-modal controllable music generation, and due to the inherent differences between modalities, music generation solutions designed for single modality are difficult to directly migrate to other modalities. For example, music generation models trained for videos are difficult to directly migrate to text, humming, etc. The embodiments of the present application can support multi-modal (picture, video, label, text, humming) input, and can generate high-quality music that matches it under different modal control.

[0295] (2) The generated music has a high degree of match with the control modality. The related technology adopts the technical route of matching the input modality with the music clips in the existing music library, that is, calculating the similarity based on the modal features and the music clip features, and taking the most similar music clip as the soundtrack result of the modality. This is heavily dependent on the richness of the music clips in the music library, and the generated music has a low degree of match with the input modality in terms of rhythm and emotion (rough match at the level of input modality and audio clip). The embodiment of the present application innovatively proposes a fine-grained control generation strategy of "modal content analysis-parsing to music elements-mapping to music representation sequence-control generation of matching music", and the generated music and the input modality are highly matched in terms of four types of music elements: emotion, style, rhythm, and notes.

[0296] (3) The quality of the generated music is guaranteed. The embodiment of the present application controls the quality of the output music by introducing a supervised multi-task filter. The multi-task design can guide the network to pay attention to the emotion and style information while judging the quality. Only samples that meet the quality score passing line and have the highest score will be used as the output result, thereby achieving high-quality music generation.

[0297] (4) The model control mode has strong scalability. The embodiment of the present application decouples the analysis of the multimodal input control signal from the music control generation part, that is, the training and control of the music generation model do not depend on a specific mode. The advantage of this design is that it is easy to add new modes as a "plug-in". For example, if you want to add a new mode of "human body posture", you only need to use the parser to analyze the characteristics of the mode, and parse it into specific music elements, and use the music elements as control signals for control, without having to retrain the music generator, so it has strong scalability.

[0298] The following is a description of an exemplary structure of a music file generation device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules in the music file generation device 455 stored in the memory 450 may include: a data acquisition module 4551, configured to acquire at least one modal data, wherein each modality is associated with at least one music element; a feature extraction module 4552, configured to perform feature extraction processing on the at least one modal data to obtain at least one modal feature; the feature extraction module 4552 is further configured to perform mapping processing on each modal feature according to the type of the music element associated with each modal feature, to obtain the element feature of each music element associated with the modal feature; a decoding module 4553, configured to map each element feature into a music representation sequence, wherein the music representation sequence includes sound information corresponding to different moments in the music file; the decoding module 4553 is further configured to perform decoding processing on the music representation sequence to obtain a music file.

[0299] In some embodiments, the modality includes an image modality; the type of music element associated with the image modality is an emotional element; the feature extraction module 4552 is configured to perform emotional type prediction processing on the modal features of the image modality to obtain a first prediction probability that the modal features belong to different emotional types; and the emotional type corresponding to the highest first prediction probability is used as the element feature of the emotional element.

[0300] In some embodiments, the modality includes a text modality; the type of the music element associated with the text modality is an emotional element; the feature extraction module 4552 is configured to determine the element features of the emotional element of the text modality by any of the following methods:

[0301] Perform emotion type prediction processing on the modal features of the text modality to obtain a second prediction probability that the modal features belong to different emotion types, and use the emotion type corresponding to the highest second prediction probability as the element feature of the emotion element; obtain the first similarities between the text features of multiple reference emotion texts and the modal features of the text modality, and use the emotion type to which the reference emotion text corresponding to the highest first similarity belongs as the element feature of the emotion element, wherein each of the reference emotion texts belongs to a different emotion type.

[0302] In some embodiments, the modality includes a label modality; the types of music elements associated with the label modality include: emotional elements and style elements; the feature extraction module 4552 is configured to obtain a mapping relationship table between labels and music elements, wherein the mapping relationship table includes: mapping relationships between different labels and different types of emotional elements, and mapping relationships between different labels and different types of style elements; according to the mapping relationship table, the modal features of the label modality are mapped to element features of at least one of the emotional elements and the style elements.

[0303] In some embodiments, the modality includes a video modality; the types of music elements associated with the video modality include: rhythmic elements and emotional elements; the modal features of the video modality include: video frame sequence, scene switching rate of the video frame sequence and average optical flow intensity; the feature extraction module 4552 is configured to determine the music speed based on the scene switching rate and the tangent function, wherein the music speed represents the number of beats per minute in the music file; the ratio of the average optical flow intensity to the video size is used as the optical flow amplitude value of the video, and the ratio between the optical flow amplitude value and the number of music frames is used as the note density; the instantaneous optical flow change rate of the video frame sequence is used as the beat intensity; the music speed, the note density and the beat intensity are combined as the element feature of the rhythmic element; the video frame sequence is processed for emotional type prediction to obtain a third prediction probability that the modal feature belongs to different emotional types; the emotional type corresponding to the highest third prediction probability is used as the element feature of the emotional element.

[0304] In some embodiments, the modality includes an audio modality; the types of music elements associated with the audio modality include: rhythmic elements and note elements; the feature extraction module 4552 is configured to transcribe the modal features of the audio modality to obtain element features of the note elements, wherein the element features of the note elements are represented by the first note sequence, and the first note sequence includes multiple notes and the playback order corresponding to each of the notes; feature extraction processing is performed on the first note sequence to obtain element features of the rhythmic elements, wherein the element features of the rhythmic elements include music speed, note density and beat intensity.

[0305] In some embodiments, the types of music elements include: note elements, style elements, emotional elements and rhythmic elements; the decoding module 4553 is configured to generate a second note sequence of at least one audio track based on the element characteristics of the note elements, wherein the second note sequence includes the playback order of each note; parameter configuration is performed on each of the second note sequences based on the rhythmic elements to obtain a third note sequence; each of the third note sequences is labeled based on the element characteristics of the emotional elements and the style elements, and the labeled third note sequences are superimposed to obtain the music representation sequence, wherein the music representation sequence includes sound information at each playback moment, and the sound information includes: style type, emotional type, notes, and note density and beat intensity corresponding to the notes.

[0306] In some embodiments, the music representation sequence includes sound information at each moment; the sound information is represented by a word-gram matrix composed of different information; the decoding module 4553 is configured to perform aggregation processing on the word-gram matrix of the sound information of the music representation sequence to obtain a word-gram encoding vector; perform multiple rounds of sound information prediction processing based on the word-gram encoding vector of the music representation sequence to obtain an updated music representation sequence, wherein the updated music representation sequence is composed of the music representation sequence and multiple predicted sound information obtained by prediction; the updated music representation sequence is format-converted to obtain a music file.

[0307] In some embodiments, the decoding module 4553 is configured to perform linear mapping processing on the word unit matrix of each of the sound information in the music representation sequence to obtain a coding embedding vector; connect each of the coding embedding vectors according to time sequence to obtain a concatenated feature vector; perform linear mapping on the concatenated feature vector, and perform position coding processing on the result of the linear mapping to obtain a word unit coding vector.

[0308] In some embodiments, the decoding module 4553 is configured to use the music representation sequence as the first music representation sequence; predict the predicted sound information of the next moment corresponding to the first music representation sequence based on the word unit encoding vector of the current music representation sequence; add the predicted sound information to the end of the current music representation sequence to obtain a second music representation sequence; use the second music representation sequence as the first music representation sequence, and perform the processing of predicting the predicted sound information of the next moment corresponding to the first music representation sequence based on the word unit encoding vector; in response to the length of the second music representation sequence reaching a preconfigured length, use the second music representation sequence as an updated music representation sequence.

[0309] In some embodiments, the number of the music files is multiple; the music file generation device further includes a screening module, the screening module being configured to perform the following processing on each of the music files after decoding the music representation sequence to obtain the music file:

[0310] The sound information of each moment of the music representation sequence is predicted and processed to obtain the prediction probability of the music representation sequence corresponding to different evaluation dimensions, wherein the evaluation dimensions include: music quality, music emotion and music style; weighted summation is performed on each of the prediction probabilities to obtain the quality index of the music file; and the music file corresponding to the highest quality index is taken as the optimal music file.

[0311] In some embodiments, the sound information is represented by a word-gram matrix; the screening module is configured to perform encoding processing on the word-gram matrix of each of the sound information to obtain the encoding features of each of the sound information; perform global average pooling processing on each of the encoding features according to the time dimension to obtain global features; normalize the global features for each of the evaluation dimensions to obtain the predicted probability of the music representation sequence corresponding to each of the evaluation dimensions.

[0312] In some embodiments, the method for generating the music file is implemented by a music generation model, which includes an encoder and a decoder; wherein the encoder is used to obtain the following features: modal features and element features of each of the music elements associated with the modal features; the decoder is used to obtain a music representation sequence and decode the music representation sequence.

[0313] In some embodiments, the music generation model further includes a filter, and when a plurality of the generated music files are present, the filter is used to determine an optimal music file from the plurality of the music files.

[0314] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer program or the computer executable instruction from the computer-readable storage medium, and the processor executes the computer program or the computer executable instruction, so that the electronic device executes the method for generating a music file described in the embodiment of the present application.

[0315] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the method for generating a music file provided by the embodiment of the present application, for example, Figure 3A A method for generating a music file is shown.

[0316] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0317] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0318] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).

[0319] As an example, the executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.

[0320] In summary, through the embodiments of the present application, the modal data is mapped to the element features corresponding to the music elements, the element features are mapped to the music representation sequence, and the music file is generated based on the music sequence. Multimodal controllable, supports the generation of music files through different modal data, and compared with the solution of generating music files based on single modal data in the related technology, the degree of freedom of obtaining music files is improved; the modal data is mapped to music elements, and then the music files are determined by the music elements, which improves the fine-grainedness of the modal data to generate music files. Compared with the solution of the related technology relying on music materials, it saves computing resources and improves the matching degree between the input modal data and the output music file, so that the generated music file can be more accurate.

[0321] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A method for generating a music file, characterized in that: The method comprises: Acquire at least one modality data, wherein each modality is associated with at least one music element; Performing feature extraction processing on the at least one modal data to obtain at least one modal feature; According to the type of the music element associated with each modal feature, mapping processing is performed on each modal feature to obtain element features of each music element associated with the modal feature; Mapping each of the element features into a music representation sequence, wherein the music representation sequence includes sound information corresponding to different moments in the music file; The music representation sequence is decoded to obtain a music file.

2. The method according to claim 1, characterized in that The modality includes an image modality; the type of the music element associated with the image modality is an emotional element; According to the type of the music element associated with each modal feature, mapping processing is performed on each modal feature to obtain the element feature of each music element associated with the modal feature, including: Performing emotion type prediction processing on the modal features of the image modality to obtain first prediction probabilities that the modal features belong to different emotion types; The emotion type corresponding to the highest first prediction probability is used as the element feature of the emotion element.

3. The method according to claim 1, characterized in that The modality includes a text modality; the type of the music element associated with the text modality is an emotional element; According to the type of the music element associated with each modal feature, mapping processing is performed on each modal feature to obtain the element feature of each music element associated with the modal feature, including: Determine the element features of the emotional element of the text modality by any of the following methods: Performing emotion type prediction processing on the modal features of the text modality to obtain second prediction probabilities that the modal features belong to different emotion types, and using the emotion type corresponding to the highest second prediction probability as the element feature of the emotion element; The first similarities corresponding to the text features of multiple reference emotion texts and the modal features of the text modality are obtained, and the emotion type to which the reference emotion text corresponding to the highest first similarity belongs is used as the element feature of the emotion element, wherein each of the reference emotion texts belongs to a different emotion type.

4. The method according to claim 1, characterized in that: The modality includes a tag modality; the types of music elements associated with the tag modality include: emotional elements and musical style elements; According to the type of the music element associated with each modal feature, mapping processing is performed on each modal feature to obtain the element feature of each music element associated with the modal feature, including: Obtaining a mapping relationship table between tags and music elements, wherein the mapping relationship table includes: mapping relationships between different tags and different types of emotion elements, and mapping relationships between different tags and different types of music style elements; According to the mapping relationship table, the modal features of the label modality are mapped to element features of at least one of the emotional elements and the musical style elements.

5. The method according to claim 1, characterized in that The modality includes a video modality; the types of music elements associated with the video modality include: rhythmic elements and emotional elements; the modality features of the video modality include: video frame sequence, scene switching rate of the video frame sequence, and average optical flow intensity; According to the type of the music element associated with each modal feature, mapping processing is performed on each modal feature to obtain the element feature of each music element associated with the modal feature, including: Determine the music speed based on the scene switching rate and the tangent function, wherein the music speed represents the number of beats per minute in the music file; The ratio of the average optical flow intensity to the video size is used as the optical flow amplitude value of the video, and the ratio of the optical flow amplitude value to the number of music frames is used as the note density; Taking the instantaneous optical flow change rate of the video frame sequence as the beat intensity; Combining the music speed, the note density and the beat intensity as element features of the rhythmic element; Performing emotion type prediction processing on the video frame sequence to obtain third prediction probabilities that the modal features belong to different emotion types; The emotion type corresponding to the highest third prediction probability is used as the element feature of the emotion element.

6. The method according to claim 1, characterized in that The modality includes an audio modality; the types of music elements associated with the audio modality include: rhythm elements and note elements; According to the type of the music element associated with each modal feature, mapping processing is performed on each modal feature to obtain the element feature of each music element associated with the modal feature, including: Performing transcription processing on the modal features of the audio modality to obtain the element features of the note elements, wherein the element features of the note elements are represented by the first note sequence, and the first note sequence includes a plurality of notes and a playback order corresponding to each of the notes; The first note sequence is subjected to feature extraction processing to obtain element features of the rhythmic elements, wherein the element features of the rhythmic elements include music speed, note density and beat intensity.

7. The method according to claim 1, characterized in that The types of music elements include: note elements, style elements, emotional elements and rhythmic elements; Mapping each of the element features into a music representation sequence includes: generating a second note sequence of at least one audio track based on the element features of the note elements, wherein the second note sequence includes a playback order of each note; Perform parameter configuration on each of the second note sequences based on the rhythmic elements to obtain a third note sequence; Each of the third note sequences is labeled based on the element features of the emotional elements and the musical style elements, and the labeled third note sequences are superimposed to obtain the music representation sequence, wherein the music representation sequence includes sound information at each playback moment, and the sound information includes: musical style type, emotional type, notes, and note density and beat intensity corresponding to the notes.

8. The method according to claim 1, characterized in that The music representation sequence includes the sound information at each moment; The sound information is represented by a word-unit matrix composed of different information; The decoding process of the music representation sequence to obtain a music file includes: Aggregating the word unit matrix of the sound information of the music representation sequence to obtain a word unit encoding vector; Performing multiple rounds of sound information prediction processing based on the word unit encoding vector of the music representation sequence to obtain an updated music representation sequence, wherein the updated music representation sequence is composed of the music representation sequence and a plurality of predicted sound information obtained by prediction; The updated music representation sequence is format-converted to obtain a music file.

9. The method according to claim 8, characterized in that The aggregating process of the word unit matrix of the sound information of the music representation sequence to obtain the word unit encoding vector includes: Performing linear mapping processing on the word-unit matrix of each of the sound information in the music representation sequence to obtain a coding embedding vector; Connecting each of the encoding embedding vectors in time order to obtain a concatenated feature vector; Linear mapping is performed on the concatenated feature vector, and position encoding is performed on the result of the linear mapping to obtain a word unit encoding vector.

10. The method according to claim 8, characterized in that The word unit encoding vector based on the music representation sequence performs multiple rounds of sound information prediction processing to obtain an updated music representation sequence, including: Using the music representation sequence as a first music representation sequence; Based on the word unit encoding vector of the current music representation sequence, predict the predicted sound information at the next moment corresponding to the first music representation sequence; Adding the predicted sound information to the end of the current music representation sequence to obtain a second music representation sequence; The second music representation sequence is used as the first music representation sequence, and the process of predicting the predicted sound information at the next moment corresponding to the first music representation sequence is performed based on the word unit encoding vector; In response to the length of the second music representation sequence reaching a preconfigured length, the second music representation sequence is used as an updated music representation sequence.

11. The method according to any one of claims 1 to 10, characterized in that: The number of the music files is multiple; After decoding the music representation sequence to obtain the music file, the method further includes: The following processing is performed on each of the music files: Predicting the sound information of each moment of the music representation sequence to obtain the prediction probability of the music representation sequence corresponding to different evaluation dimensions, wherein the evaluation dimensions include: music quality, music emotion and music style; Performing weighted summation on each of the predicted probabilities to obtain a quality index of the music file; The music file corresponding to the highest quality index is regarded as the optimal music file.

12. The method according to claim 11, characterized in that The sound information is represented by a word unit matrix; The predicting process of the sound information of each moment of the music representation sequence to obtain the prediction probability of the music representation sequence corresponding to different evaluation dimensions includes: Performing encoding processing on the word-unit matrix of each of the sound information to obtain encoding features of each of the sound information; Performing global average pooling processing on each of the encoding features according to the time dimension to obtain a global feature; The global feature is normalized for each evaluation dimension to obtain the prediction probability of the music representation sequence corresponding to each evaluation dimension.

13. The method according to any one of claims 1 to 9, characterized in that: The method for generating the music file is implemented by a music generation model, which includes an encoder and a decoder; wherein the encoder is used to obtain the following features: modal features and element features of each of the music elements associated with the modal features; the decoder is used to obtain a music representation sequence and decode the music representation sequence.

14. The method according to claim 12, characterized in that The music generation model further includes a filter, and when a plurality of the generated music files are present, the filter is used to determine an optimal music file from the plurality of the music files.

15. A device for generating a music file, characterized in that: The device comprises: A data acquisition module configured to acquire at least one modality of data, wherein each modality is associated with at least one musical element; A feature extraction module, configured to perform feature extraction processing on the at least one modal data to obtain at least one modal feature; The feature extraction module is further configured to perform mapping processing on each of the modal features according to the type of the music element associated with each of the modal features, so as to obtain element features of each music element associated with the modal feature; A decoding module configured to map each of the element features into a music representation sequence, wherein the music representation sequence includes sound information corresponding to different moments in the music file; The decoding module is further configured to decode the music representation sequence to obtain a music file.

16. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; The processor is used to implement the method for generating a music file as described in any one of claims 1 to 14 when executing the computer executable instructions or computer program stored in the memory.

17. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the method for generating a music file according to any one of claims 1 to 14 is implemented.

18. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the method for generating a music file according to any one of claims 1 to 14 is implemented.