Music generation method, music generation device, electronic equipment and storage medium

Through artificial intelligence, accompaniment audio and singing audio are generated and audio synthesis is performed, which solves the problem of time and cost of traditional music creation, and achieves efficient personalization of music creation and improves user experience.

CN120279868APending Publication Date: 2025-07-08XIAOMI EV TECH CO LTD +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510519153.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The traditional music creation model is cumbersome, time-consuming and costly, and has set a high creative threshold, making it difficult to meet the needs of non-professional creators.

Method used

Artificial intelligence technology is used to generate accompaniment audio and singing audio respectively, and target music is generated through audio synthesis processing, improving the adaptability of accompaniment audio and singing audio, and promoting the two to achieve close integration in multiple dimensions such as rhythm, tone, and emotional expression.

Benefits of technology

It greatly improves the efficiency and user experience of music creation, lowers the technical threshold, and makes the music creation process more efficient and personalized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279868A_ABST
    Figure CN120279868A_ABST
Patent Text Reader

Abstract

The invention provides a music generation method, a music generation device, electronic equipment and a storage medium, and the music generation method comprises the steps: after receiving a music generation instruction, generating an accompaniment audio and a singing audio according to the music generation instruction, carrying out the audio synthesis of the accompaniment audio and the singing audio, and generating and outputting target music. Therefore, when the method is applied to the vehicle intelligent cabin scene, the accompaniment audio and the singing audio are respectively generated by adopting an artificial intelligence technology, and the accompaniment audio and the singing audio are synthesized to obtain the target music, so that the adaptation degree of the accompaniment audio and the singing audio in music content creation can be effectively improved; the accompaniment audio and the singing audio are closely fused in multiple dimensions such as rhythm, tone and emotional expression, so that the user experience is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a music generation method, a music generation device, an electronic device, and a storage medium. Background Art

[0002] Traditional music generation has long relied on manual music arrangement, instrument performance, and professional recording processes. This creative mode has many significant limitations. For example, this mode has a cumbersome process, resulting in a long time-consuming traditional music creation. From the initial creative concept to the completion of the final work, it often takes weeks or even months; moreover, the cost is high, involving expenses in aspects such as instrument purchase, venue rental, and personnel remuneration. This undoubtedly sets a relatively high threshold for non-professional creators. Summary of the Invention

[0003] This application aims to at least partly solve one of the technical problems in the related art.

[0004] To this end, this application proposes a music generation method, a device, an electronic device, and a storage medium. This application uses artificial intelligence technology to separately generate an accompaniment audio and a singing audio, and synthesizes the accompaniment audio and the singing audio to obtain a target music, thereby being able to effectively improve the adaptation degree of the accompaniment audio and the singing audio in music content creation, promoting their close integration in multiple dimensions such as rhythm, pitch, and emotional expression, and thus greatly improving the user experience.

[0005] The first aspect embodiment of this application proposes a music generation method, including:

[0006] In response to a received music generation instruction, separately generate an accompaniment audio and a singing audio;

[0007] Perform audio synthesis processing on the accompaniment audio and the singing audio, and generate and output a target music.

[0008] The second aspect embodiment of this application proposes a music generation device, including a module for executing the above music generation method, where the module includes:

[0009] A generation module, configured to separately generate an accompaniment audio and a singing audio in response to a received music generation instruction;

[0010] A synthesis module, configured to perform audio synthesis processing on the accompaniment audio and the singing audio, and generate and output a target music.

[0011] The third aspect embodiment of this application proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above music generation method are implemented.

[0012] In the fourth aspect embodiment of the present application, a non - transitory computer - readable storage medium is proposed, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the above - mentioned music generation method are realized.

[0013] In the fifth aspect embodiment of the present application, a computer program product is proposed, on which a computer program is stored. When the program is executed by a processor, the steps of the above - mentioned music generation method are realized.

[0014] For the music generation method, music generation device, electronic device and storage medium proposed in the present application, after receiving a music generation instruction, according to the music generation instruction, an accompaniment audio and a singing audio are respectively generated, and then the accompaniment audio and the singing audio are subjected to audio synthesis processing to generate and output a target music. Thus, when this method is applied to the vehicle intelligent cockpit scenario, artificial intelligence technology is used to respectively generate an accompaniment audio and a singing audio, and the accompaniment audio and the singing audio are synthesized to obtain a target music. Thereby, it can effectively improve the adaptation degree between the accompaniment audio and the singing audio in music content creation, promote their close integration in multiple dimensions such as rhythm, pitch, and emotional expression, and further greatly improve the user experience.

[0015] Additional aspects and advantages of the present application will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above - mentioned and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0017] Figure 1 is a schematic flow chart of a music generation method provided by an embodiment of the present application;

[0018] Figure 2 is a schematic flow chart of an accompaniment audio generation provided by an embodiment of the present application;

[0019] Figure 3 is a schematic flow chart of generating an accompaniment audio based on a control parameter set provided by an embodiment of the present application;

[0020] Figure 4 is a schematic structural diagram of a diffusion model provided by an embodiment of the present application;

[0021] Figure 5 is a schematic flow chart of a singing audio generation provided by an embodiment of the present application;

[0022] Figure 6 is a schematic flow chart of generating a singing audio based on a control parameter set provided by an embodiment of the present application;

[0023] Figure 7 Schematic diagram of a music generation device provided by an embodiment of the present application;

[0024] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific implementation manners

[0025] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0026] The music generation method, music generation device, electronic device, and storage medium of the embodiments of the present application will be described below with reference to the accompanying drawings.

[0027] Figure 1 Flow chart of a music generation method provided by an embodiment of the present application.

[0028] It should be noted that the music generation method of the embodiments of the present application can be applied to a music generation device. In some possible embodiments, the music generation device can be configured in an electronic device so that the electronic device can execute the music generation function. Additionally, in some possible embodiments, the music generation device can also be software in the electronic device, etc.

[0029] Among them, the electronic device can be a device equipped with a music generation device, including but not limited to: vehicles, servers, terminals, etc. For example, the terminal can be a device such as a mobile phone, a speaker, etc. The terminal can also be referred to as a terminal device (terminal), user equipment (UE for short), mobile station (MS for short), mobile terminal device (MT for short), etc.

[0030] The forms of the terminal are diverse, including automobiles with communication functions, intelligent vehicles, mobile phones, wearable devices, tablets (Pads), computers with wireless transceiver functions, virtual reality (VR for short) terminals, augmented reality (AR for short) terminals, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grid, wireless terminals in transportation safety, wireless terminals in smart city, wireless terminals in smart home, and so on.

[0031] The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal.

[0032] The music generation method proposed in the embodiments of the present application can cover the vehicle field in its application scenarios, and is particularly applicable to the specific carrier of the intelligent cockpit of an automobile. Taking the intelligent karaoke system configured in the cockpit as an example, this method can realize diversified music generation functions to meet the entertainment needs of users in the cockpit environment.

[0033] As Figure 1 shown, the music generation method of the embodiments of the present application includes:

[0034] Step S101, in response to receiving a music generation instruction, generate an accompaniment audio and a singing audio respectively.

[0035] The process of receiving the music generation instruction is as follows:

[0036] After the user starts the music generation software, clicks or wakes up by voice through the main interface of the software to enter the "music generation" function. Immediately after the operation, the interface jumps, and an interactive instruction input window pops up. In the window, the user can freely fill in or input the music generation instruction by voice through the text input box. For example, clearly express the requirement as: "Generate a lively pop song with the theme of 'love' and the rhythm set to 120 BPM".

[0037] The process of generating the accompaniment audio is as follows:

[0038] When the software receives a music generation instruction input by the user through the interaction interface, it first conducts a structured parsing process on the music generation instruction. This process includes the following steps: First, through natural language processing technology, extract the core parameters in the music generation instruction, including style parameters (determined as "pop song"), rhythm parameters (determined as 120 BPM), duration parameters (implied or explicitly specified), and theme parameters (such as "love"); Second, based on the parsed style parameters, the software calls the built-in instrument configuration library to automatically match the standard instrument combination template corresponding to the pop song style, which presets the timbre parameters and playing rules of basic instruments such as guitars, basses, and drums; Third, based on the parsed rhythm parameters (120 BPM), the software calls the rhythm generation module to generate a basic rhythm skeleton according to the rhythm characteristics of pop music, which defines the drumbeat distribution, the rhythm pattern of the bass line, and the chord progression mode of the guitar; Fourth, combining the rhythm skeleton and style parameters, the software generates an upbeat accompaniment melody through a melody generation algorithm, which integrates common chord progressions (such as I-V-vi-IV) and melody motives (such as jumping intervals, repeated rhythm patterns) in pop music; Fifth, according to the theme parameter "love", the software calls the sound effect enhancement module to dynamically insert romantic-style sound effect elements into the accompaniment, such as superimposing the glissando of the strings in the chorus part to enhance the emotional tension, and adding the tremolo and appoggiatura of the piano at the melody transition to improve the musical fineness; Finally, the software mixes the generated instrument timbres, rhythms, melodies, and sound effect elements in multiple tracks to generate an accompaniment audio, and supports export to standard audio formats such as WAV and MP3. This can significantly improve the automation level of the accompaniment audio generation in the music creation process and lower the technical threshold for users.

[0039] The generation process of the singing audio is as follows:

[0040] After the software receives the music generation instruction input by the user through the interaction interface, it first performs a structured parsing process on the instruction to extract the core theme parameters contained in the instruction (such as "love"). Subsequently, it calls the pre-trained lyric generation model to automatically generate lyric content that conforms to semantic logic and emotional expression based on the theme parameters. For example, "In this romantic time, our love blossoms like a flower, every moment is filled with sweetness, and being with you is my most beautiful dream". After the lyric generation is completed, the software matches the vocal line characteristics corresponding to the target style according to the preset singing style library (such as pop, rock, folk, etc.). For example, for the singing requirements of pop songs, it selects the sweet timbre of a female voice. The timbre parameters include the pitch range (such as C3 - C5), formant characteristics (such as enhanced high-frequency overtones), and pronunciation methods (such as a mixture of breathy and true voices). Next, the software uses a speech synthesis engine to map and transform the generated lyric text and vocal line characteristics to generate the initial singing audio. At the same time, the software calls the preset melody template library according to the target style, extracts melody segments that match the emotion and rhythm of the lyrics, and aligns the lyrics and melody through the Dynamic Time Warping (DTW) algorithm to ensure the temporal consistency of syllables and notes, generating the singing audio. After that, post-processing optimization can be performed on the generated singing audio, including reverb enhancement, panning adjustment, and dynamic range compression, and the finished audio is saved in standard formats such as WAV or MP3. This can significantly improve the automation level of singing audio generation in music content creation and reduce the technical threshold for users.

[0041] Step S102, perform audio synthesis processing on the accompaniment audio and the singing audio, and output the target music.

[0042] After the software generates the accompaniment audio and the singing audio, it loads the previously generated accompaniment audio and singing audio files from the preset storage path and starts the automatic verification process. This automatic verification process covers format compatibility detection (such as the sampling rate being uniformly 44.1kHz and the bit depth matching being 16 bits) and quality preprocessing (such as peak level normalization and silent segment clipping) to ensure that there are no format conflicts and sound quality defects in the audio materials. Subsequently, the rhythm characteristics of the two audio segments are extracted through a beat detection algorithm, and elastic alignment is performed on the segments with minor time deviations based on the dynamic time warping algorithm to achieve the rhythm synchronization of singing and accompaniment.

[0043] In the audio synthesis stage, the software uses a multi-track mixing engine for intelligent fusion: first, dynamically adjust the sound pressure level ratio of vocals and accompaniment according to the preset music style template (such as pop, rock) (for example, the vocals in pop songs are usually 3-6dB higher than the accompaniment); second, optimize the frequency band distribution through the multi-band equalizer (EQ) (such as improving the clarity of the mid-frequency of the vocals and enhancing the strength of the low frequency of the accompaniment), and apply dynamic range compression to improve the overall loudness; finally, combine the reverberation algorithm to simulate the sense of space (such as room reverberation, hall reverberation) to make the music more layered. During the audio synthesis process, monitor the spectrum balance in real time to avoid auditory fatigue caused by frequency band conflicts.

[0044] After the audio is synthesized, the software automatically saves the target music file in the user-specified format (such as WAV, MP3) and generates a standard file containing metadata such as song title, artist, and duration. At the same time, a prompt window pops up on the software interface, providing interactive buttons such as "Play", "Share", and "Edit". Among them, after the user clicks the play button, the software loads and plays the music in streaming form; click the share button to upload the work with one click through a third-party application platform. In addition, the software also supports users to retroactively adjust the synthesis parameters (such as rebalancing the vocal volume and changing the reverb type) to achieve flexible iteration of the creative process.

[0045] It should be noted that in terms of the practical application and expansion of music generation technology, the music generation method proposed in this application shows a strong scene adaptation ability and can be widely used in the field of diversified music creation. For example, for audio-visual content such as movies, TV series, animations and games, the software can automatically generate background music with high fit based on factors such as plot rhythm, scene atmosphere (such as suspense, cheerfulness, epic) and character emotions, and enhance narrative immersion; in the production of commercial advertisements and promotional videos, this method can combine brand tonality, product characteristics and target audience preferences to quickly customize background music with emotional resonance and improve communication effects; for individual users, the software supports the generation of unique and exclusive music works based on user-entered keywords (such as "birthday", "graduation" and "travel"), emotional labels (such as "romantic", "passionate" and "healing") or specific music styles (such as jazz, electronic, folk songs) to meet personalized expression needs. Through cross-scenario technology migration, this method not only lowers the professional threshold for music creation, but also provides content producers with efficient and intelligent creation tools to help the deep integration of the music industry and multiple industries.

[0046] Thus, in the music generation method of the embodiments of the present application, when the software receives a music generation instruction input by the user, it can automatically generate a complete target music. The music generation method of the embodiments of the present application can effectively improve the adaptation degree between the accompaniment audio and the singing audio in music content creation, promote their close integration in multiple dimensions such as rhythm, pitch, and emotional expression, and thus greatly improve the user experience.

[0047] Figure 2 It is a schematic flowchart of the generation of an accompaniment audio provided by the embodiments of the present application.

[0048] As Figure 2 shown, the generation process of the accompaniment audio in the embodiments of the present application includes:

[0049] Step S201, in response to receiving the music generation instruction, parsing and processing the music generation instruction to obtain a set of control parameters.

[0050] After the user starts the music creation software, a music generation instruction can be input in the form of natural language through the software interface. For example, the user inputs the instruction: "Please generate a song in jazz style, with a slow rhythm, highlighting the saxophone performance, and the lyrics theme centered around 'Encounter in the Night City'." After the software receives this instruction, it immediately transmits it to the built-in instruction parsing module.

[0051] The instruction parsing module uses natural language processing technology to structurally parse the instruction input by the user, extract key creative elements, and map them to a set of control parameters recognizable by the software. Among them, the set of control parameters includes: a style parameter, used to determine the style characteristics of the target music; a prompt parameter, used to specify the performance elements of the target music.

[0052] The software identifies the style characteristic description in the instruction (such as "jazz style") through keyword matching and semantic analysis, and converts it into a predefined style code. For example, "jazz style" is mapped to the style parameter JZ001, and this code is associated with characteristic parameters such as the typical chord progression, rhythm pattern, and instrument configuration of jazz.

[0053] The software extracts the creative prompt information in the instruction (such as "slow rhythm", "saxophone performance"), and converts it into a structured prompt parameter:

[0054] Rhythm parameter: Parse "slow rhythm" as rhythm_01, corresponding to a slow rhythm template with a BPM range of 60 - 80;

[0055] Instrument parameter: Parse "saxophone performance" as instrument_01, specifying the saxophone as the main instrument, and associating its tone library and performance technique parameters.

[0056] The software integrates the parsed style parameters and hint parameters into a control parameter set including style parameters (JZ001) and hint parameters (rhythm_01, instrument_01).

[0057] Step S202: Generate an accompaniment audio based on the control parameter set.

[0058] As an implementable way of step S202, it includes the following steps:

[0059] Step S2021: Determine the tempo information and chord information based on the style parameters and hint parameters.

[0060] According to the style parameters (JZ001, representing the jazz style) and hint parameters (rhythm_01, indicating a slow rhythm; instrument_01, specifying the saxophone), generate the tempo information and chord information through the following steps.

[0061] The software retrieves the tempo value that matches the style parameters and hint parameters from the preset stylized tempo database. For example, for the slow rhythm of the jazz style, the software preferentially selects the tempo within the range of 60 - 80 BPM, and further filters based on the weight of the hint parameters (such as the intensity of the "slow" keyword), and finally determines the tempo information as 80 BPM.

[0062] The software combines the style parameters with the jazz chord rule library to generate the adapted chord information (i.e., chord sequence). For example, based on the JZ001 style parameters, the software calls the II-V-I progression (such as Dm7 - G7 - Cmaj7) or its variants (such as Cmaj7 - Am7 - Dm7 - G7), and dynamically adjusts the complexity and color of the chord information (such as adding chord extensions or substitute chords) according to the instrument parameters (such as saxophone) and emotional tendency (such as slow) in the hint parameters.

[0063] Step S2022: Encode the chord information based on the tempo information to obtain the first encoded feature.

[0064] After obtaining the tempo information (80 BPM) and chord information (Cmaj7 - Am7 - Dm7 - G7), the software combines the diffusion model framework as shown in Figure 4 and uses a conditional encoder to encode the chord information (i.e., the input conditional information of the conditional encoder is the chord information), and the process is as follows:

[0065] The software first maps the chord names to discrete indices and converts them into vector representations of a fixed dimension through one-hot encoding or an embedding layer. For example, Cmaj7 corresponds to [1, 0, 0, 0], and Am7 corresponds to [0, 1, 0, 0]. Subsequently, the tempo information is quantized into discrete categories or directly input as continuous numerical values, and the tempo is used as a context constraint through a conditional encoder to dynamically adjust the encoding weights of the chord vectors. The software inputs the chord vector sequence into the forward process of the diffusion model, gradually adding noise to learn the data distribution; in the reverse denoising process, the tempo information is input as a condition to guide the model to generate a chord encoding representation that conforms to the tempo constraint. Finally, the software processes the chord sequence in chronological order, generates the conditional encoding vector for each chord, and integrates it into a sequence-level feature representation. For example, the first encoding feature of the chord sequence Cmaj7-Am7-Dm7-G7 at a tempo of 80 BPM may be: [[1, 0, 0, 0, tempo feature sub-vector], [0, 1, 0, 0, tempo feature sub-vector], [0, 0, 1, 0, tempo feature sub-vector], [0, 0, 0, 1, tempo feature sub-vector]].

[0066] Among them, the conditional encoder includes but is not limited to VAE (Variational Autoencoder), one-hot encoder, or a combination thereof.

[0067] In step S2023, the prompt parameters are encoded to obtain the second encoding feature.

[0068] After obtaining the prompt parameters (such as tempo_01 (slow tempo) and instrument_01 (saxophone)), the software follows the process as Figure 4 shown, and uses a text encoder to convert the prompt parameters into vector representations of a fixed dimension (that is, the text information input to the text encoder is the prompt parameters), and the process is as follows:

[0069] The software first converts the prompt parameters into a standardized format and processes them through word segmentation or sub-word segmentation. Subsequently, a pre-trained language model (such as BERT, GPT) or a custom embedding layer is used to map the text sequence to a dense vector. For example, the encoded vectors of tempo_01 (slow tempo) and instrument_01 (saxophone) may be [0.12, -0.34,...] and [0.56, 0.78,...] respectively.

[0070] The second encoding feature is the concatenation result of the encoded vectors of each prompt parameter: [0.12, -0.34,...] + [0.56, 0.78,...] = [0.12, -0.34,... 0.56, 0.78,...].

[0071] In step S2024, the accompaniment audio is generated based on the first encoding feature and the second encoding feature.

[0072] After obtaining the first encoded feature and the second encoded feature, the software first concatenates the first encoded feature and the noise feature to obtain a first concatenated feature, then uses a cross-attention mechanism and a diffusion mechanism to interact the second encoded feature and the first concatenated feature to obtain a first interaction feature, and then decodes the first interaction feature to obtain the accompaniment audio.

[0073] For example, the first encoded feature is a vector of length 100, and the noise feature is a vector of length 50. After concatenation, a first concatenated feature of length 150 is obtained. Then, the cross-attention mechanism is used to calculate the degree of association between each element in the second encoded feature and each element in the first concatenated feature, and the diffusion mechanism is used to weight and fuse the features according to these degrees of association to obtain the first interaction feature. Finally, the first interaction feature is decoded and converted into an audio signal to generate the accompaniment audio. Among them, the decoding process can use a neural network model, such as a recurrent neural network (RNN) or a long short-term memory network (LSTM), to gradually convert the feature vector into an audio sample.

[0074] Thus, the music generation method of the present application, after receiving the music generation instruction, parses and processes the music generation instruction to obtain a control parameter set, and automatically generates the accompaniment audio according to the control parameter set.

[0075] Figure 5 It is a schematic flowchart of the generation of a singing audio provided by an embodiment of the present application.

[0076] As Figure 5 shown, the singing audio generation process of the embodiment of the present application includes:

[0077] Step S501, in response to receiving the music generation instruction, parse and process the music generation instruction to obtain a control parameter set.

[0078] After the user starts the music creation software, the music generation instruction can be input in the form of natural language through the software interface. For example, the user inputs the instruction: "Please generate a jazz-style song with a slow rhythm, highlighting the saxophone performance, and the lyrics theme is centered around 'Encounter in the City at Night'." After the software receives this instruction, it immediately transmits it to the built-in instruction parsing module.

[0079] The instruction parsing module uses natural language processing technology to structurally parse the user input instruction, extract key creation elements and map them to a control parameter set recognizable by the software. Among them, the control parameter set includes: a style parameter for determining the style characteristics of the target music; a prompt parameter for specifying the performance elements of the target music; and a lyrics parameter for providing the lyrics content of the target music.

[0080] Among them, it should be noted that the extraction processes of the style parameters and the prompt parameters in the control parameter set have been described in detail in step S201 of the above embodiments, and will not be elaborated here. The following describes the method for obtaining the lyrics parameters.

[0081] The software parses the descriptive word "Encounter in the City at Night" related to the target music in the music generation instruction, and calls a preset lyrics generation model or algorithm according to this descriptive word to generate the lyrics content as the lyrics parameters. For example, the generated lyrics may be "The city lights are dim at night, I happened to meet you on the street, and the melody of the saxophone spreads slowly like our story."

[0082] In addition, if the music generation instruction input by the user contains lyrics content, the software can directly obtain the lyrics content input by the user as the lyrics parameters.

[0083] The software integrates the parsed style parameters and prompt parameters into a control parameter set including style parameters (JZ001), prompt parameters (Tempo_01, Instrument_01), and lyrics parameters ("The city lights are dim at night, I happened to meet you on the street, and the melody of the saxophone spreads slowly like our story").

[0084] Step S502, generate a singing audio based on the control parameter set.

[0085] As an implementation process of step S502, it includes the following steps:

[0086] Step S5021, generate melody information corresponding to the lyrics parameters according to the lyrics parameters, tempo information, and chord information; where the tempo information and chord information are determined based on the style parameters and prompt parameters.

[0087] Among them, it should be noted that the process of determining the tempo information and chord information based on the style parameters and prompt parameters is as detailed in the description of step S2021 in the above embodiments, and will not be elaborated here specifically.

[0088] After the software obtains the lyrics parameters, tempo information (such as 80BPM), and chord information (such as Cmaj7 - Am7 - Dm7 - G7), it generates melody information corresponding to the lyrics parameters through an intelligent melody generation algorithm. The process is as follows:

[0089] The software parses the lyrics parameters, extracts their prosodic features (such as level and oblique tones, rhyming) and emotional tendencies (such as lyrical, exciting), and then combines the tempo information and chord information, and uses a conditional generation model for melody design. The conditional generation model maps each word or phrase in the lyrics to a sequence of notes to ensure high - pitch matching, rhythm adaptation, and style enhancement.

[0090] For example, for the lyrics "The city lights are dim at night", the software generates the following melody information:

[0091] Pitch: Based on the key of C major, the melody starts from the dominant G (corresponding to "night") and gradually rises to the tonic C (corresponding to "city"), forming an emotional progression;

[0092] Rhythm: 4 / 4 time is used, the first two words "night" correspond to eighth notes, and the last two words "city" are expanded to dotted quarter notes, creating a soothing atmosphere;

[0093] Other features: Tremolo decoration is added at the "Lanshan" part, and a muted piano tone is used to echo the hazy mood of the lyrics.

[0094] Step S5022, decomposing the lyrics parameters at the phoneme level to obtain phoneme information.

[0095] The software performs factor-level operations on the input lyrics parameters and converts them into continuous phoneme information (i.e., factor sequence). The specific process is as follows:

[0096] The software first standardizes the parameters of the lyrics, including removing punctuation marks and unifying character encoding, to ensure the standardization of the input text. Then, a predefined phoneme dictionary or phoneme annotation model (such as the G2P model based on deep learning) is used to map each Chinese character or word in the lyrics to a corresponding phoneme sequence. For example, the Chinese character "夜" is mapped to the phoneme "yè"; the Chinese character "晚" is mapped to the phoneme "wǎn"; and so on. The complete lyrics "夜城市灯火呆珊" are decomposed into a phoneme sequence: "yèwǎn de chéng shìdēng huǒlán shān". Finally, the generated phoneme sequence is stored in the form of a string or vector, and each phoneme corresponds to an independent encoding unit for subsequent association with music features (such as pitch, rhythm). For example, the phoneme "yè" can be represented as [phoneme ID_1], the phoneme "wǎn" can be represented as [phoneme ID_2], and so on.

[0097] Example

[0098] Take the lyrics "The city lights are dim at night" as an example:

[0099] Original lyrics: "The city lights are dim at night."

[0100] Decomposition result: Phoneme information = "yèwǎn de chéng shìdēng huǒlán shān".

[0101] Application scenario: The phoneme information will be used as the basic input for generating melody information, and coordinated with information such as tempo and chords to ensure that the pronunciation rhythm of the lyrics accurately matches the rhythm of the music.

[0102] Step S5023: Based on the tempo information, encode the phoneme information and the melody information to obtain the third encoded feature.

[0103] After obtaining the tempo information (80 BPM), the phoneme information, and the melody information, the software combines with the diffusion model framework as shown, and uses a conditional encoder to encode the phoneme information and the melody information (that is, the input conditional information of the conditional encoder is the phoneme information and the melody information). The process is as follows: Figure 4 As shown, for the phoneme information, the one-hot encoding method is adopted to map each phoneme into a vector representation of a fixed dimension. For example, if the phoneme dictionary contains N phonemes, each phoneme is encoded as a vector of length N, where the value at the corresponding position is 1 and the rest are 0.

[0104] For the phoneme information, the one-hot encoding method is used to map each phoneme into a vector representation of a fixed dimension. For example, if the phoneme dictionary contains N phonemes, each phoneme is encoded as a vector of length N, where the value at the corresponding position is 1 and the rest are 0.

[0105] The melody information is multi-dimensionally encoded based on features such as pitch, duration, and dynamics:

[0106] Pitch: Convert the note to a MIDI pitch value (0 - 127) and normalize it to the range [0, 1].

[0107] Duration: Calculate the relative duration ratio according to the note duration (such as quarter note, eighth note) and encode it as a floating-point number.

[0108] Dynamics: Extract the dynamic range of the note (such as forte-piano markings) and quantize it into discrete levels (such as 0 - 10 levels).

[0109] The above features are combined into a melody feature vector through a fully connected layer or a convolutional layer.

[0110] Convert the tempo information (such as 80 BPM) to a scalar value and map it to a feature vector with the same dimension as the phoneme and melody features through a linear transformation. Then concatenate the phoneme encoding vector, the melody feature vector, and the tempo feature vector along the feature dimension to form the third encoded feature.

[0111] For example, assume the input is:

[0112] Phoneme sequence: ["yè", "wǎn", "de"] → One-hot encoded as [1, 0,..., 0], [0, 1,..., 0], [0, 0, 1,..., 0] (assuming N = 100);

[0113] Melody segment: [pitch = 60 (C4), duration = 0.5 (eighth note), dynamics = 7] → Encoded as [0.469, 0.5, 0.7];

[0114] Tempo = 80 BPM → Encoded as [0.8] (assuming linear mapping to [0, 1]).

[0115] Finally, the third encoded feature is as follows:

[0116] [Phoneme one-hot encoding 1, Phoneme one-hot encoding 2,..., Phoneme one-hot encoding N, Melody pitch, Melody duration, Melody intensity, Tempo].

[0117] Step S5024: Generate a singing audio according to the third encoded feature and the second encoded feature, where the second encoded feature is obtained by encoding the prompt parameters.

[0118] It should be noted that the process of encoding the prompt parameters by the second encoded feature is as detailed in the above embodiment for step S2023, and will not be elaborated here specifically.

[0119] After the software obtains the second encoded feature and the third encoded feature, it first concatenates the third encoded feature and the noise feature to obtain a second concatenated feature, then uses a cross-attention mechanism and a diffusion mechanism to interact the second encoded feature and the second concatenated feature to obtain a second interaction feature, and then decodes the second interaction feature to obtain a singing audio.

[0120] For example, the third encoded feature is a vector with a length of 100, and the noise feature is a vector with a length of 50. After concatenation, a second concatenated feature with a length of 150 is obtained. Then, the cross-attention mechanism is used to calculate the correlation degree between each element in the second encoded feature and each element in the second concatenated feature, and the diffusion mechanism is used to weight and fuse the features according to these correlation degrees to obtain a second interaction feature. Finally, the second interaction feature is decoded to convert it into an audio signal to generate a singing audio. Among them, the decoding process can use a neural network model, such as a recurrent neural network (RNN) or a long short-term memory network (LSTM), to gradually convert the feature vector into an audio sample.

[0121] Thus, the music generation method of the present application, after receiving the music generation instruction, parses and processes the music generation instruction to obtain a control parameter set, and automatically generates a singing audio according to the control parameter set.

[0122] To sum up, the music generation method of the embodiment of the present application, after receiving a music generation instruction, generates an accompaniment audio and a singing audio respectively according to the music generation instruction, and then performs audio synthesis processing on the accompaniment audio and the singing audio to generate and output a target music. Thus, this method uses artificial intelligence technology to generate an accompaniment audio and a singing audio respectively, and synthesizes the accompaniment audio and the singing audio to obtain a target music, thereby effectively improving the adaptation degree between the accompaniment audio and the singing audio in music content creation, promoting their close integration in multiple dimensions such as rhythm, pitch, and emotional expression, and thus greatly improving the user experience.

[0123] To implement the above embodiments, an embodiment of the present application further provides a music generation device.

[0124] Figure 7 It is a schematic structural diagram of a music generation device provided by an embodiment of the present application.

[0125] As Figure 7 shown, the music generation device 700 of the embodiment of the present application is used for the module that executes the above music generation method, wherein the module includes: a generation module 710 and a synthesis module 720.

[0126] Among them, the generation module 710 is used to respectively generate an accompaniment audio and a singing audio in response to a received music generation instruction;

[0127] The synthesis module 720 is used to perform audio synthesis processing on the accompaniment audio and the singing audio, and generate and output a target music.

[0128] In an embodiment of the present application, when the generation module 710 is used to respectively generate an accompaniment audio and a singing audio, it includes:

[0129] Parse and process the music generation instruction to obtain a set of control parameters;

[0130] Based on the set of control parameters, respectively generate an accompaniment audio and a singing audio.

[0131] In an embodiment of the present application, the set of control parameters includes:

[0132] A style parameter, used to determine the style characteristics of the target music;

[0133] A prompt parameter, used to specify the performance elements of the target music.

[0134] In an embodiment of the present application, when the generation module 710 is used to generate an accompaniment audio based on the set of control parameters, it includes:

[0135] Based on the style parameter and the prompt parameter, determine the tempo information and chord information;

[0136] Based on the tempo information, encode the chord information to obtain a first encoded feature;

[0137] Encode the prompt parameter to obtain a second encoded feature;

[0138] Generate an accompaniment audio according to the first encoded feature and the second encoded feature.

[0139] In an embodiment of the present application, when the generation module 710 is used to generate an accompaniment audio according to the first encoded feature and the second encoded feature, it includes:

[0140] Concatenate the first encoded feature and the noise feature to obtain the first concatenated feature;

[0141] Adopt a cross-attention mechanism and a diffusion mechanism to interact the second encoded feature and the first concatenated feature to obtain the first interaction feature;

[0142] Decode the first interaction feature to obtain the accompaniment audio.

[0143] In an embodiment of the present application, the control parameter set further includes:

[0144] Lyric parameters for providing the lyric content of the target music;

[0145] When the generation module 710 is used to generate the singing audio based on the control parameter set, it includes:

[0146] Generate the melody information corresponding to the lyric parameters according to the lyric parameters, the tempo information and the chord information; wherein, the tempo information and the chord information are determined based on the style parameters and the hint parameters;

[0147] Decompose the lyric parameters at the phoneme level to obtain the phoneme information;

[0148] Encode the phoneme information and the melody information based on the tempo information to obtain the third encoded feature;

[0149] Generate the singing audio according to the third encoded feature and the second encoded feature; wherein, the second encoded feature is obtained by encoding the hint parameters.

[0150] In an embodiment of the present application, when the generation module 710 is used to generate the singing audio according to the third encoded feature and the second encoded feature, it includes:

[0151] Concatenate the third encoded feature and the noise feature to obtain the second concatenated feature;

[0152] Adopt a cross-attention mechanism and a diffusion mechanism to interact the second encoded feature and the second concatenated feature to obtain the second interaction feature;

[0153] Decode the second interaction feature to obtain the singing audio.

[0154] In an embodiment of the present application, the method for obtaining the lyric parameters includes at least one of the following:

[0155] Receive the lyric content input by the user through the music generation instruction;

[0156] Parse the description words related to the target music in the music generation instruction and generate the lyric content according to the description words.

[0157] It should be noted that the foregoing explanatory description of the embodiments of the music generation method also applies to the music generation device of this embodiment, and will not be elaborated here.

[0158] After receiving a music generation instruction, the music generation device according to the embodiments of the present application generates an accompaniment audio and a singing audio respectively through a generation module according to the music generation instruction, and performs audio synthesis processing on the accompaniment audio and the singing audio through a synthesis module to generate and output a target music. Thus, the device uses artificial intelligence technology to generate an accompaniment audio and a singing audio respectively, and synthesizes the accompaniment audio and the singing audio to obtain the target music, thereby effectively improving the adaptation degree of the accompaniment audio and the singing audio in music content creation, promoting their close integration in multiple dimensions such as rhythm, pitch, and emotional expression, and further significantly improving the user experience.

[0159] To implement the above embodiments, the present application also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the music generation method described in any of the foregoing embodiments are implemented.

[0160] Figure 8 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0161] Referring to Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0162] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0163] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0164] The power component 806 provides power to the various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0165] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0166] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0167] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0168] The sensor assembly 814 includes one or more sensors for providing status assessment of various aspects for the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0169] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as Wi-Fi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0170] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the steps of the above method.

[0171] To implement the above embodiments, the present application also proposes a non-transitory computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the music generation method described in any of the foregoing method embodiments are implemented.

[0172] To implement the above embodiments, the present application also proposes a computer program product having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the music generation method described in any of the foregoing method embodiments are implemented.

[0173] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0174] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0175] Any process or method description depicted in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application pertain.

[0176] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0177] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0178] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0179] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately physically for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0180] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A music generation method, characterized in that, Including: In response to receiving a music generation instruction, respectively generate an accompaniment audio and a singing audio; Perform audio synthesis processing on the accompaniment audio and the singing audio, and generate and output a target music.

2. The method according to claim 1, characterized in that, The respectively generating an accompaniment audio and a singing audio includes: Perform parsing processing on the music generation instruction to obtain a set of control parameters; Based on the set of control parameters, respectively generate the accompaniment audio and the singing audio.

3. The method according to claim 2, characterized in that, The set of control parameters includes: A style parameter, used to determine the style characteristics of the target music; A hint parameter, used to specify the performance elements of the target music.

4. The method according to claim 3, characterized in that, Based on the set of control parameters, generating the accompaniment audio includes: Based on the style parameter and the hint parameter, determine the tempo information and chord information; Based on the tempo information, encode the chord information to obtain a first encoded feature; Encode the hint parameter to obtain a second encoded feature; According to the first encoded feature and the second encoded feature, generate the accompaniment audio.

5. The method according to claim 4, wherein The generating the accompaniment audio according to the first encoded feature and the second encoded feature includes: Concatenate the first encoded feature and a noise feature to obtain a first concatenated feature; Adopt a cross-attention mechanism and a diffusion mechanism to interact the second encoded feature and the first concatenated feature to obtain a first interaction feature; Decode the first interaction feature to obtain the accompaniment audio.

6. The method according to claim 3, wherein The set of control parameters further includes: A lyrics parameter, used to provide the lyrics content of the target music; Based on the set of control parameters, generating the singing audio includes: According to the lyrics parameter, tempo information and chord information, generate melody information corresponding to the lyrics parameter; wherein, the tempo information and chord information are determined based on the style parameter and the hint parameter; Perform phoneme-level decomposition on the lyrics parameter to obtain phoneme information; Based on the tempo information, encode the phoneme information and the melody information to obtain a third encoded feature; According to the third encoded feature and the second encoded feature, generate the singing audio; wherein, the second encoded feature is obtained by encoding the hint parameter.

7. The method according to claim 6, characterized in that, The generating the singing audio according to the third encoded feature and the second encoded feature includes: Concatenate the third encoded feature and a noise feature to obtain a second concatenated feature; Adopt a cross-attention mechanism and a diffusion mechanism to interact the second encoded feature and the second concatenated feature to obtain a second interaction feature; Decode the second interaction feature to obtain the singing audio.

8. The method according to claim 6, wherein The ways of obtaining the lyrics parameter include at least one of the following: Receive the lyrics content input by the user through the music generation instruction; Parse the description words related to the target music in the music generation instruction, and generate the lyrics content according to the description words.

9. A music generating device, characterized in that, Including a module for executing the music generation method according to any one of claims 1-8, wherein the module includes: A generation module, used to respectively generate an accompaniment audio and a singing audio in response to the received music generation instruction; A synthesis module for performing audio synthesis processing on the accompaniment audio and the singing audio to generate and output target music.

10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the music generation method according to any one of claims 1-8.

11. A non-transitory computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the music generation method according to any one of claims 1-8 are implemented.

12. A computer program product, characterized in that, It includes a computer program that, when executed by the processor, implements the steps of the music generation method according to any one of claims 1-8.

Citation Information

Cited By

  • Blind music creation method and electronic equipment

    CN120954363A

  • Blind music creation method and electronic device

    CN120954363B