Music synthesis method, device and equipment, storage medium and vehicle
By using a lightweight music source separation model in an embedded system to separate and edit music source signals, the problem of real-time generation of instrument tracks in existing technologies is solved, enabling convenience and efficiency for improvisation and live performance.
Patent Information
- Application Number
- CN202410693726.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-02
AI Technical Summary
In existing music production technologies, music synthesis relies on cloud services or high-performance personal computers, which makes it impossible to generate instrument tracks in real time, improvise or live performance, and results in low efficiency in music synthesis.
The system employs a lightweight music source separation model embedded in the system to separate music source signals, identify and edit instrument track frames, and generate target music through merging processing, including modification, addition, and deletion operations, supporting real-time creation and live performance.
It does not rely on cloud services or high-performance computers, is highly convenient, can generate target music in real time, improves the efficiency of music synthesis, and facilitates improvisation and live performance.
Smart Images

Figure CN121053931A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, device, storage medium, and vehicle for music synthesis. Background Technology
[0002] Current music production technology is undergoing a wave of technological innovation, especially in the area of Music Source Separation (MSS). MSS technology plays a crucial role in music production, as it breaks down mixed music signals into individual instrument tracks (Stems).
[0003] In existing technologies, music players typically require preprocessing of the music source to synthesize instrument tracks, which are then edited before the music is synthesized. This music synthesis relies on cloud services or high-performance personal computers, resulting in low convenience. Furthermore, the preprocessed music source cannot generate instrument tracks in real time, hindering real-time creation or modification and leading to low efficiency in music synthesis, thus limiting improvisation and live performances. Summary of the Invention
[0004] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides a music synthesis method, apparatus, device, storage medium and vehicle to improve the convenience of improvisation or live performance and improve the efficiency of synthesized music.
[0005] In a first aspect, embodiments of this disclosure provide a music synthesis method, including:
[0006] Acquire a music source signal, wherein the music source signal contains at least two instrument tracks;
[0007] The music source signal is separated by a music source separation model to obtain multiple instrument tracks of the music source signal. The music source separation model is a lightweight model for embedded systems.
[0008] For each instrument track, identify the target track frame of the instrument track, and perform editing operations on the target track frame to obtain the target instrument track;
[0009] Multiple target instrument tracks are merged to export the target music.
[0010] In some embodiments, the editing operation includes: modification operation, addition operation, and deletion operation, wherein the modification operation includes at least one of the following: modifying volume, modifying timbre, modifying rhythm, and modifying effect; the addition operation includes at least one of the following: expanding audio track, adding loop playback count; and the deletion operation includes at least one of the following: trimming audio track, reducing loop playback count.
[0011] The editing operation on the target audio track frame includes:
[0012] Analyze user needs for the target music;
[0013] Based on the requirements, add or delete the target audio track frame to obtain candidate instrument audio tracks;
[0014] Based on the aforementioned requirements, the candidate instrument tracks are modified to obtain the target instrument track.
[0015] In some embodiments, the music source signal includes multiple music source signal segments, each music source signal segment being composed of a preset frame of music source signal; the method further includes:
[0016] Acquire a real-time transmitted music source signal segment, wherein the music source signal segment carries a timestamp;
[0017] The music source signal segment is separated by a music source separation model to obtain multiple instrument track segments of the music source signal segment;
[0018] For each instrument track segment, identify the target track frame of the instrument track segment, and perform editing operations on the target track frame to obtain the target instrument track segment;
[0019] By merging multiple target instrument track fragments, a target music fragment is obtained;
[0020] The target music is obtained by splicing the target music segments according to the timestamps.
[0021] In some embodiments, acquiring the music source signal includes:
[0022] Acquire a music source signal in a preset format, and store the music source signal in a preset space;
[0023] Accordingly, the process of merging multiple target instrument tracks to export the target music includes:
[0024] The multiple target instrument tracks are stored separately in the preset space;
[0025] In response to the user selecting the target music, the audio of the multiple target instrument tracks is played simultaneously.
[0026] In some embodiments, the step of splicing the target music segment according to the timestamp to obtain the target music includes:
[0027] The target music segments are sorted according to the timestamps to obtain sorted target music segments;
[0028] Process overlapping or blank segments of the sorted target music segments, and then splice the sorted target music segments in sequence to obtain the target music.
[0029] In some embodiments, the method includes constructing a music source separation model, which includes a depthwise separable convolutional model, a knowledge distillation model, and an extraction model;
[0030] The step of separating the music source signal using a music source separation model to obtain multiple instrument tracks of the music source signal includes:
[0031] The original depthwise separable convolutional model is used as the student model in the knowledge distillation model. The trained teacher model is then imitated using knowledge distillation techniques to obtain the depthwise separable convolutional model.
[0032] The music source signal is extracted from the music source signal using the depthwise separable convolution model, and the music source signal is initially separated based on the music source signal to generate multiple original music source signals;
[0033] The extraction model is used to optimize the multiple original sound source signals to obtain the instrument track corresponding to each music source signal.
[0034] Secondly, embodiments of this disclosure provide a music synthesis apparatus, comprising:
[0035] The acquisition module is used to acquire a music source signal, wherein the music source signal contains at least two instrument tracks;
[0036] The separation module is used to separate the music source signal using a music source separation model to obtain multiple instrument tracks of the music source signal. The music source separation model is a lightweight model of an embedded system.
[0037] The editing module is used to identify the target track frame of each instrument track, perform editing operations on the target track frame, and obtain the target instrument track.
[0038] The processing module is used to merge multiple target instrument tracks and export the target music.
[0039] Thirdly, embodiments of this disclosure provide a portable electronic device, including:
[0040] Memory;
[0041] Processor; and
[0042] Computer programs;
[0043] The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in the first aspect.
[0044] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the method described in the first aspect.
[0045] Fifthly, embodiments of this disclosure also provide a vehicle, including: a music synthesis device as described in the second aspect; or a portable electronic device as described in the third aspect; or a computer-readable storage medium as described in the fourth aspect.
[0046] The music synthesis method, apparatus, device, storage medium, and vehicle provided in this disclosure acquire a music source signal containing at least two instrument tracks; separate the music source signal using a music source separation model to obtain multiple instrument tracks, where the music source separation model is a lightweight model of an embedded system; for each instrument track, identify the target track frame, edit the target track frame to obtain the target instrument track; merge the multiple target instrument tracks to export the target music. Compared to existing technologies, this method does not rely on cloud services or high-performance personal computers, offering high convenience. It allows editing of the multiple instrument tracks obtained from the music source signal separation to generate the target instrument track, thereby obtaining the target music, improving the efficiency of music synthesis and facilitating improvisation or live performances. Attached Figure Description
[0047] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0048] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 This is a flowchart of a music synthesis method provided in an embodiment of the present disclosure;
[0050] Figure 2 A schematic diagram illustrating an application scenario provided by an embodiment of this disclosure;
[0051] Figure 3 This is a schematic diagram of the structure of a music source signal provided in an embodiment of the present disclosure;
[0052] Figure 4This is a schematic diagram of the structure of a music source signal provided in an embodiment of the present disclosure;
[0053] Figure 5 A schematic diagram illustrating an application scenario provided by an embodiment of this disclosure;
[0054] Figure 6 This is a schematic diagram of the structure of the music synthesis device provided in the embodiments of this disclosure;
[0055] Figure 7 A schematic diagram of the structure of a portable electronic device provided in an embodiment of this disclosure;
[0056] Figure 8 A schematic diagram of the structure of a portable electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0057] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0058] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0059] This disclosure provides a music synthesis method, which will be described below with reference to specific embodiments.
[0060] Figure 1 This is a flowchart illustrating a music synthesis method provided in an embodiment of this disclosure. The method can be executed by a music synthesis device, which can be implemented using software and / or hardware. This device can be configured in an electronic device that is easily portable, such as a terminal, specifically including a Walkman, mobile phone, tablet computer, or smartwatch. Furthermore, this method can be applied to music synthesis, improvisation, and live performance scenarios. It is understood that the music synthesis method provided in this embodiment can also be applied to other scenarios.
[0061] The following is about Figure 1 The music synthesis method shown is introduced below, and the specific steps of this method are as follows:
[0062] S101. Obtain a music source signal, wherein the music source signal contains at least two instrument tracks.
[0063] Instrument tracks are one of the fundamental elements in music production. They refer to dividing a music recording into several independent tracks, each containing different instruments or sound elements. These tracks can be vocals, synthesizers, drum kits, bass, violins, pianos, etc., and each is an independent audio file.
[0064] Acquire a music source signal, which may be obtained from a recording studio, memory card, live performance, or music platform, and the music source signal contains at least two instrument tracks.
[0065] S102. The music source signal is separated by a music source separation model to obtain multiple instrument tracks of the music source signal. The music source separation model is a lightweight model of an embedded system, which can achieve efficient music source signal separation.
[0066] By separating the music source signal using a trained music source separation model, multiple instrument tracks of the music source signal are obtained. It can be understood that this music source separation model is a lightweight model for embedded systems.
[0067] It is understandable that the music source separation model can be a combination of a deep separable convolutional model, a knowledge distillation model, and an extraction model. Specifically, it can be any one or more of sequential combination, parallel combination, and result integration. Sequential combination involves first using a deep separable convolutional model for initial separation, then inputting the initial separation result into a knowledge distillation model for optimization, and finally using an extraction model for precise separation and optimization. Parallel combination involves inputting the music source signal into the deep separable convolutional model, the knowledge distillation model, and the extraction model respectively, and then fusing or selectively using the outputs of these three models. Result integration involves combining the outputs of the deep separable convolutional model, the knowledge distillation model, and the extraction model through post-processing methods (such as weighted averaging, voting, etc.) to obtain multiple instrument tracks. Specifically, the extraction model can be an f_0 extraction model used to extract instrument tracks and convert them into Musical Instrument Digital Interface (MIDI) signals for mixing between multiple devices.
[0068] Extracting the fundamental frequency of audio signals and converting it into a MIDI signal using the f_0 extraction model is a common task in Music Information Retrieval (MIR) and audio processing. MIDI signals allow music devices to exchange information such as notes, volume, and rhythm. Specific steps include: preprocessing, fundamental frequency extraction, note mapping, and MIDI generation.
[0069] Specifically, preprocessing the music source signal: such as noise reduction and normalization of the music source signal.
[0070] Fundamental frequency extraction: Audio analysis algorithms are used to estimate the fundamental frequency of an audio signal. These algorithms can be based on time-domain (e.g., autocorrelation function) or frequency-domain (e.g., Fourier transform and wavelet transform) algorithms. The music source signal is transformed from the time domain to the frequency domain, the individual frequency components in the music source signal are identified, and the fundamental frequency, i.e., the lowest frequency component, is determined by analyzing the spectrogram.
[0071] Note mapping: Mapping the fundamental frequency estimate to musical notes involves converting the fundamental frequency into MIDI key codes, based on scales and modes in music theory. For example, the frequency of a standard A4 note is 440Hz, and its corresponding MIDI key code is 69. The frequencies of other notes are integer or half-integer multiples of this frequency, and the corresponding MIDI key codes change accordingly. The formula is: midinumber = 69 + 12 × log2(f_0 / 440).
[0072] Generate MIDI: Generate MIDI based on mapped musical notes. This includes determining the start and end times of the musical notes, as well as dynamic information such as volume and tempo. The generated MIDI events are exported as standard MIDI file formats, which can be read and used by any MIDI-enabled device or software.
[0073] The exchange of MIDI signals between multiple devices promotes diversity in collaborative creation and performance. Specifically, terminal 1 and terminal 2 exchange MIDI signals for logic control, or terminal 3 participates in the performance as a slave device.
[0074] S103. For each instrument track, identify the target track frame of the instrument track, and perform editing operations on the target track frame to obtain the target instrument track.
[0075] For each instrument track, in response to the user's creative intent or needs, the target track frame of the instrument track is identified, and the instrument track is edited through interactive operations such as clicking, double-clicking, and swiping to obtain the target instrument track.
[0076] S104. Merge multiple target instrument tracks to export the target music.
[0077] This function merges multiple target instrument tracks and exports the target music file, which can be in MP3, WAV, or FLAC format.
[0078] Specifically, such as Figure 2As shown, the target instrument track f0 is mixed using a Musical Instrument Digital Interface (MIDI). 1 and f0 2 The signals are converted into MIDI signals, and the MIDI mixer controls the transmission of the MIDI signals to the synthesizer. The synthesizer then performs merging processing (also known as synthesis processing) on the MIDI signals to output the target music.
[0079] This embodiment of the disclosure acquires a music source signal, which contains at least two instrument tracks; it separates the music source signal using a music source separation model to obtain multiple instrument tracks, where the music source separation model is a lightweight model of an embedded system; for each instrument track, it identifies the target track frame, edits the target track frame to obtain the target instrument track; and it merges multiple target instrument tracks to export the target music. Compared to existing technologies, this method does not rely on cloud services or high-performance personal computers, is highly convenient, and allows editing of multiple instrument tracks obtained from the music source signal to generate the target instrument track, thereby obtaining the target music. This improves the efficiency of music synthesis and facilitates improvisation or live performances.
[0080] Based on the above embodiments, the editing operations include: modification operations, addition operations, and deletion operations. The modification operations include at least one of the following: modifying volume, modifying timbre, modifying rhythm, and modifying effects. The addition operations include at least one of the following: expanding the audio track and adding loop playback counts. The deletion operations include at least one of the following: trimming the audio track and reducing the loop playback counts.
[0081] Specifically, volume can be modified using professional audio editing software such as Adobe Audition, Audacity, or Logic Pro X; or by creating automatic control clips to automate volume modification, such as controlling the volume of specific instruments in a playlist to achieve volume gradients or other complex volume changes; or by using dedicated volume modification software such as Audio Processing Master or SoundVolumeView to quickly modify the volume of audio files.
[0082] Specifically, you can modify the timbre using music effects. You can refer to other people's timbre and add effects such as reverb, distortion, and equalizer adjustments.
[0083] Specifically, modifying rhythm involves altering the temporal arrangement of notes in an audio signal, including the start time, duration, and intervals of the notes. This can be achieved using quantization tools, slicing and reconstruction tools, and time stretching / compression algorithms.
[0084] Specifically, modifying effects means adding audio effects, including fade-in / fade-out, echo, reverb, etc., which can be directly applied to the audio track. In particular, by adjusting the fade-in / fade-out effect, the beginning and end of the audio can be made smoother and more natural.
[0085] Specifically, extending a track means copying an existing instrument track and pasting it into the target location, or using the extension function to lengthen the audio.
[0086] Specifically, trimming an audio track typically involves using selection tools and the cut function to mark the start and end points of the audio region to be trimmed, and then performing the trimming operation.
[0087] Specifically, the loop count is achieved by setting the loop start and end points, as well as the number of loops. Correspondingly, increasing the loop count means increasing the number of loops at the start and end points, while decreasing the loop count means decreasing the number of loops at the start and end points.
[0088] The editing operation of the target audio track frame includes: analyzing the user's needs for the target music; adding or deleting the target audio track frame based on the needs to obtain candidate instrument tracks; and modifying the candidate instrument tracks based on the needs to obtain the target instrument track.
[0089] like Figure 3 As shown, 31, 32, 33, and 34 are various instrument tracks obtained from the separation of music source signals. Analyzing the user's needs for the target music, 35 is the target track frame on instrument track 31. The target track frame is the track frame that needs to be modified by the user. According to the user's needs, the target track frame is added or deleted to obtain candidate instrument tracks. According to the user's needs, the candidate instrument tracks are modified to obtain the target instrument track.
[0090] Specifically, the modification operations also include time stretching and shifting. Time stretching changes the playback speed of the audio track frame to lengthen or shorten the length of the audio segment; shifting moves the audio track frame forward or backward on the timeline to change the position of the audio segment.
[0091] Specifically, volume adjustment includes gain adjustment and balance control. Gain adjustment refers to increasing or decreasing the overall volume of the audio track; balance control refers to adjusting the volume of different frequency bands, such as bass, midrange and treble.
[0092] Specifically, modifying the timbre includes using an equalizer to adjust the frequency distribution of the track, changing the brightness or darkness of the timbre; changing specific frequency bands of the timbre through high-pass, low-pass, or band-pass filters; and adding distortion or saturation effects to make the timbre more personalized.
[0093] Specifically, the modification effects include reverb, delay, and modulation effects. Reverb refers to adding a room or hall reverb effect to the audio track to simulate different acoustic environments; delay refers to adding echo or delay effects to create a sense of space or special musical effects; modulation effects refer to using effects such as vibrato and phase shift to add dynamic changes to the timbre.
[0094] This embodiment of the disclosure analyzes the user's needs for target music; adds or deletes target audio tracks based on the needs to obtain candidate instrument tracks; modifies candidate instrument tracks based on the needs to obtain target instrument tracks; and modifies target audio tracks in real time according to user needs, thereby improving the convenience of improvisation.
[0095] In some embodiments, the music source signal includes multiple music source signal segments, each music source signal segment being composed of a preset frame music source signal. The method further includes: acquiring music source signal segments transmitted in real time, each music source signal segment carrying a timestamp; separating the music source signal segments using a music source separation model to obtain multiple instrument track segments of the music source signal segments; for each instrument track segment, identifying a target track frame of the instrument track segment, performing an editing operation on the target track frame to obtain a target instrument track segment; merging multiple target instrument track segments to obtain a target music segment; and splicing the target music segments according to the timestamps to obtain the target music.
[0096] Specifically, such as Figure 4As shown, the music source signal is 41, which includes 5 music source signal segments: music source signal segment A, music source signal segment B, music source signal segment C, music source signal segment D, and music source signal segment E. Each music source signal segment is composed of preset frames of music source signals, and the preset frames can be any positive integer between 1 and N. The music source signal segments transmitted in real time are obtained. Each music source signal segment carries a timestamp, meaning the music source signal is processed in frames. Assuming each frame is m in size, Overlap-save (OLS) and... The music source signal segment A is accumulated to a preset frame using the overlap-add (OLA) method. Then, the music source signal segment A is separated using a music source separation model to obtain multiple instrument track segments. For example, if there are three instrument track segments, they can be denoted as instrument track segment A1, instrument track segment A2, and instrument track segment A3, respectively. For each instrument track segment, the target track frame is identified, and editing operations are performed on the target track frame to obtain the target instrument track segment. Specifically, for instrument track segment A1, the target track frame a1 of instrument track segment A1 is identified, the user's needs for the target music are analyzed, and editing operations (i.e., adding and / or deleting operations, and modification operations) are performed on the target track frame a1 according to the user's needs. The process involves several steps: First, for instrument track segment A2, the target instrument track frame a2 is identified. Then, the user's needs for the target music are analyzed, and editing operations (addition, deletion, and modification) are performed on the target track frame a2 according to the user's needs, resulting in target instrument track segment A′2. Similarly, for instrument track segment A3, the target track frame a3 is identified. The user's needs for the target music are analyzed, and editing operations (addition, deletion, and modification) are performed on the target track frame a3 according to the user's needs, resulting in target instrument track segment A′3. Finally, multiple target instrument track segments are merged: target instrument track segments A′1, A′2, and A′3 are merged to obtain target music segment A′. During the processing of music source signal segment A, music source signal segment B can be acquired simultaneously in real time. The implementation principle and process of music source signal segment B are consistent, and target music segment B′ can be obtained through music source signal segment B. This embodiment will not be described in detail. According to the timestamp of each frame of the music source signal, the target music segment C′, target music segment D′ and target music segment E′ are obtained sequentially through music source signal segment C, music source signal segment D and music source signal segment E. According to the timestamp of the music source signal segments, the target music segment A′, target music segment B′, target music segment C′, target music segment D′ and target music segment E′ are spliced together to obtain the target music.
[0097] It is understood that, in other embodiments, the target instrument track fragments A′1, B′1, C′1, D′1, and E′1 can be spliced together to obtain target instrument track 1; the target instrument track fragments A′2, B′2, C′2, D′2, and E′2 can be spliced together to obtain target instrument track 2; the target instrument track fragments A′3, B′3, C′3, D′3, and E′3 can be spliced together to obtain target instrument track 3; and target instrument track 1, 2, and 3 can be merged to obtain the target music.
[0098] This disclosure describes a music source signal comprising multiple music source signal segments, each composed of preset frame music source signals. The process involves acquiring real-time transmitted music source signal segments; separating these segments using a music source separation model to obtain multiple instrument track segments; identifying the target track frame for each instrument track segment; editing the target track frame to obtain the target instrument track segment; merging multiple target instrument track segments to obtain the target music segment; and concatenating the target music segments according to timestamps to obtain the target music. This method, without requiring cloud services or high-performance personal computers, generates the target music in real-time by segmenting the music source signal, further improving the efficiency of music synthesis.
[0099] In some embodiments, acquiring the music source signal includes: acquiring a music source signal in a preset format, the music source signal being stored in a preset space; correspondingly, merging multiple target instrument tracks to export target music includes: storing the multiple target instrument tracks separately in the preset space; and simultaneously playing the audio of the multiple target instrument tracks in response to a user selecting the target music.
[0100] A music source signal in a preset format is acquired and stored in a preset space, which can be a memory card, etc., and the preset format can be MP3, WAV, or FLAC. When the terminal is powered on, the maximum segment length is calculated, and the music source signal is separated using a music source separation model to obtain multiple instrument tracks. For each instrument track, the target track frame is identified, and the maximum segment length of the target track frame is edited to obtain the target instrument track. The multiple target instrument tracks are stored separately in the preset space to obtain the target music. When the user selects the target music for playback, the terminal synchronously plays the audio of the multiple target instrument tracks so that the user can hear the target music.
[0101] Optionally, the step of splicing the target music segments according to the timestamp to obtain the target music includes: sorting the target music segments according to the timestamp to obtain sorted target music segments; processing overlapping or blank segments of the sorted target music segments, and splicing the sorted target music segments in sequence to obtain the target music.
[0102] Specifically, the target music segments are sorted according to their timestamps to obtain sorted target music segments. The timestamps typically represent the start and / or end times of the target music segments within the musical work. Overlapping or blank segments among multiple target music segments are identified based on the timestamps. These overlapping or blank segments are processed, and the sorted target music segments are then spliced together sequentially. During the splicing process, it is ensured that the start and end times of each target music segment match the timestamp information to maintain the continuity and accuracy of the music, resulting in the target music. The spliced target music is previewed to check for any problems. If problems are found, the splicing process is corrected and adjusted until a satisfactory result is achieved, yielding the target music.
[0103] Specifically, the system handles overlapping or blank segments of the sorted target music segments. If the timestamps of two target music segments overlap, the system can choose to ignore the overlapping parts, merge the overlapping parts, or select a specific part of one of the segments to avoid overlap. If there is a time interval between two target music segments, blanks, other music segments, or transition effects can be inserted to fill the time interval.
[0104] This embodiment of the disclosure acquires a music source signal in a preset format, stores the music source signal in a preset space, and separates the music source signal using a music source separation model to obtain multiple instrument tracks of the music source signal; for each instrument track, it identifies the target track frame of the instrument track, performs editing operations on the target track frame to obtain the target instrument track; and stores the multiple target instrument tracks separately in the preset space, so that in response to the user selecting target music, the audio of multiple target instrument tracks can be played simultaneously, thereby improving the flexibility of the music synthesis method.
[0105] In some embodiments, the method includes constructing a music source separation model, which includes a depthwise separable convolutional model, a knowledge distillation model, and an extraction model;
[0106] The deep separable convolutional model mainly includes input processing, feature extraction, separation, and reconstructed output. Input processing: The music source signal is used as input to the deep separable convolutional model, typically a time-frequency representation (such as a Short-Time Fourier Transform (STFT)) or waveform representation. Feature extraction: Features are extracted from the music source signal through the convolutional layers of the deep separable convolutional model. These features may include information at different frequencies and time scales. Separation: Multiple output channels in the deep separable convolutional model correspond to different instrument tracks. Each output channel learns to recognize and separate features associated with a specific instrument. Reconstructed output: The results from each output channel are combined to reconstruct the independent instrument tracks.
[0107] The knowledge distillation model mainly includes a teacher model, knowledge distillation, the separation process, and the output. Teacher Model: A powerful teacher model (such as a deep neural network) is trained to perform the music source separation task. Knowledge Distillation: A smaller student model is trained using the output of the teacher model. The student model attempts to mimic the output of the teacher model using fewer computational resources. Separation Process: During training, the student model learns how to separate music source signals. The trained student model can be used independently for music source signal separation tasks. Output: The student model outputs multiple instrument tracks, which are similar to the output of the teacher model.
[0108] The extraction model primarily includes signal decomposition, instrument recognition, and track extraction. Signal decomposition: The extraction model relies on specific signal processing techniques to decompose the music source signal. For example, it may use filter-based methods or spectral analysis to separate different frequency components. Instrument recognition: The extraction model typically includes components capable of identifying the sounds of different instruments, which can be achieved through training a classifier or using techniques such as template matching. Track extraction: After identifying a specific instrument sound, the extraction model can extract individual tracks from the instrument sound.
[0109] The step of separating the music source signal using a music source separation model to obtain multiple instrument tracks of the music source signal includes: using the original depthwise separable convolutional model as a student model in a knowledge distillation model, and imitating the trained teacher model through knowledge distillation technology to obtain the depthwise separable convolutional model; extracting sound source signals from the music source signal using the depthwise separable convolutional model, initially separating the music source signal based on the sound source signal to generate multiple original sound source signals; and optimizing the multiple original sound source signals using the extraction model to obtain the instrument track corresponding to each music source signal.
[0110] The original deep separable convolutional model is used as the student model in the knowledge distillation model. Through knowledge distillation, the trained teacher model is imitated to obtain the deep separable convolutional model. The deep separable convolutional model extracts the source signals from the music source signals, and based on these, the music source signals are initially separated to generate multiple original source signals. The extraction model optimizes these multiple original source signals to obtain the instrument track corresponding to each music source signal. The deep separable convolutional model reduces computation and the number of parameters; the knowledge distillation model reduces model size and improves inference speed; and the extraction model extracts the fundamental frequency to generate music digital interface signals, enabling MIDI signal exchange between devices and improving the flexibility of music creation and production.
[0111] In other words, the music source signal is input into a music source separation model, which is typically built based on deep learning and signal processing algorithms. This model has the ability to extract and separate different sound sources from the music source signal. The music source separation model identifies the input music source signal by analyzing its spectral characteristics, temporal properties, and audio structure. After identifying multiple sound source signals, the model separates them, i.e., decomposes and reconstructs the audio signal. It extracts each sound source signal from the original music source signal while preserving their original sound quality and characteristics, resulting in multiple sound source signals corresponding to multiple instrument tracks. For example, if the original music source signal contains the sounds of guitar, drums, and bass, the music source separation model can separate the tracks corresponding to guitar, drums, and bass respectively. These tracks can exist independently for subsequent audio processing or editing.
[0112] This embodiment uses the original depthwise separable convolutional model as the student model in the knowledge distillation model. Through knowledge distillation, the trained teacher model is imitated to obtain the depthwise separable convolutional model. The depthwise separable convolutional model is used to extract sound source signals from the music source signals. Based on these sound source signals, the music source signals are initially separated to generate multiple original sound source signals. The extraction model is then used to optimize these multiple original sound source signals to obtain the instrument track corresponding to each music source signal. This clarifies the construction of the music source separation model, reduces the computational load and number of parameters, decreases the model size, improves inference speed, and further enhances the flexibility of the music synthesis method.
[0113] In some embodiments, such as Figure 5As shown, the target instrument track f0, extracted through a music source separation model, can be combined with human voice. This involves inputting both the human voice and the target instrument track f0 into a Vocoder for encoding. The combined audio undergoes real-time automatic pitch correction (AutoTune), and the target music is then output. A Vocoder is a digital signal processor used to convert analog audio into digital signals. This method allows singers to achieve accurate pitch while producing synthesized sounds that possess both human vocal quality and adherence to melodic lines. It enhances the flexibility and expressiveness of music composition, facilitating improvisational live performances and music production.
[0114] Figure 6 This is a schematic diagram of the structure of a music synthesis apparatus provided in an embodiment of this disclosure. The music synthesis apparatus can be a terminal as described in the above embodiments, or it can be a component or assembly within that terminal. The music synthesis apparatus provided in this embodiment can execute the processing flow provided in the music synthesis method embodiments, such as... Figure 6 As shown, the music synthesis device 60 includes: an acquisition module 61, a separation module 62, an editing module 63, and a processing module 64; wherein,
[0115] Acquisition module 61 is used to acquire a music source signal, wherein the music source signal contains at least two instrument tracks;
[0116] The separation module 62 is used to separate the music source signal through a music source separation model to obtain multiple instrument tracks of the music source signal. The music source separation model is a lightweight model of an embedded system.
[0117] Editing module 63 is used to identify the target track frame of each instrument track, and perform editing operations on the target track frame to obtain the target instrument track;
[0118] The processing module 64 is used to merge multiple target instrument tracks and export the target music.
[0119] Optionally, the editing operations include: modification operations, addition operations, and deletion operations. The modification operations include at least one of the following: modifying volume, modifying timbre, modifying rhythm, and modifying effects. The addition operations include at least one of the following: expanding the audio track and adding loop playback counts. The deletion operations include at least one of the following: trimming the audio track and reducing the loop playback counts. The editing module 63 is also used to analyze the user's needs for the target music; perform addition or deletion operations on the target audio track frame based on the needs to obtain candidate instrument tracks; and perform modification operations on the candidate instrument tracks based on the needs to obtain the target instrument track.
[0120] Optionally, the music source signal includes multiple music source signal segments, each segment consisting of a preset frame of music source signal.
[0121] The acquisition module 61 is also used to acquire a real-time transmitted music source signal segment, wherein the music source signal segment carries a timestamp;
[0122] The separation module 62 is also used to separate the music source signal segment through the music source separation model to obtain multiple instrument track segments of the music source signal segment;
[0123] The editing module 63 is also used to identify the target track frame of each instrument track segment, and perform editing operations on the target track frame to obtain the target instrument track segment.
[0124] The processing module 64 is also used to merge multiple target instrument track segments to obtain a target music segment; and to splice the target music segment according to the timestamp to obtain the target music.
[0125] Optionally, the acquisition module 61 is also used to acquire a music source signal in a preset format, wherein the music source signal is stored in a preset space;
[0126] Correspondingly, the processing module 64 is also used to store the multiple target instrument tracks separately in the preset space; and to play the audio of the multiple target instrument tracks simultaneously in response to the user selecting the target music.
[0127] Optionally, the processing module 64 is further configured to sort the target music segments according to the timestamp to obtain sorted target music segments; process overlapping or blank segments of the sorted target music segments, and splice the sorted target music segments in sequence to obtain the target music.
[0128] Optionally, the music synthesis device 60 further includes: a construction module for constructing a music source separation model, the music source separation model including a depthwise separable convolutional model, a knowledge distillation model, and an extraction model; a separation module 62, which is further used to use the original depthwise separable convolutional model as a student model in the knowledge distillation model, and to imitate the trained teacher model through knowledge distillation technology to obtain the depthwise separable convolutional model; to extract sound source signals from the music source signals through the depthwise separable convolutional model, and to initially separate the music source signals based on the sound source signals to generate multiple original sound source signals; and to optimize the multiple original sound source signals through the extraction model to obtain the instrument track corresponding to each music source signal.
[0129] Figure 6The music synthesis apparatus of the illustrated embodiment can be used to execute the technical solution of the music synthesis method embodiment described above. Its implementation principle and technical effect are similar, and will not be repeated here.
[0130] Figure 7 This is a schematic diagram of the structure of a portable electronic device provided in an embodiment of the present disclosure. The portable electronic device can be a terminal as described in the above embodiments. The portable electronic device provided in this embodiment of the present disclosure can execute the processing flow provided in the music synthesis method embodiments, such as… Figure 7 As shown, the portable electronic device 70 includes: a memory 71, a processor 72, a computer program, and a communication interface 73; wherein the computer program is stored in the memory 71 and configured to be executed by the processor 72 using the music synthesis method described above.
[0131] Optionally, such as Figure 8 As shown, the portable electronic device also includes: Storage (i.e., the memory 71 mentioned above), HCI (human-computer interface), MCU (microcontroller unit), Power (power system), MIDI (music digital interface), Codec (audio codec), Input (audio input), Output (audio output), Display (display screen), and Led (LED display light), wherein HCI (human-computer interface) is the user control interface, such as a touch screen or buttons.
[0132] In addition, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the music synthesis method described in the above embodiments.
[0133] Furthermore, this disclosure also provides a vehicle that includes a music synthesis device as described in the above embodiments; or a portable electronic device as described in the above embodiments; or a computer-readable storage medium as described in the above embodiments.
[0134] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0135] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for music synthesis, characterized in that, The method includes: Acquire a music source signal, wherein the music source signal contains at least two instrument tracks; The music source signal is separated by a music source separation model to obtain multiple instrument tracks of the music source signal. The music source separation model is a lightweight model for embedded systems. For each instrument track, identify the target track frame of the instrument track, and perform editing operations on the target track frame to obtain the target instrument track; Multiple target instrument tracks are merged to export the target music.
2. The method according to claim 1, characterized in that, The editing operations include: modification operations, addition operations, and deletion operations. The modification operations include at least one of the following: modifying volume, modifying timbre, modifying rhythm, and modifying effects. The addition operations include at least one of the following: expanding the audio track and adding loop playback counts. The deletion operations include at least one of the following: trimming the audio track and reducing the loop playback counts. The editing operation on the target audio track frame includes: Analyze user needs for the target music; Based on the requirements, add or delete the target audio track frame to obtain candidate instrument audio tracks; Based on the aforementioned requirements, the candidate instrument tracks are modified to obtain the target instrument track.
3. The method according to claim 1, characterized in that, The music source signal includes multiple music source signal segments, each music source signal segment being composed of a preset frame of music source signal; the method further includes: Acquire a real-time transmitted music source signal segment, wherein the music source signal segment carries a timestamp; The music source signal segment is separated by a music source separation model to obtain multiple instrument track segments of the music source signal segment; For each instrument track segment, identify the target track frame of the instrument track segment, and perform editing operations on the target track frame to obtain the target instrument track segment; By merging multiple target instrument track fragments, a target music fragment is obtained; The target music is obtained by splicing the target music segments according to the timestamps.
4. The method according to claim 1, characterized in that, The acquisition of the music source signal includes: Acquire a music source signal in a preset format, and store the music source signal in a preset space; Accordingly, the process of merging multiple target instrument tracks to export the target music includes: The multiple target instrument tracks are stored separately in the preset space; In response to the user selecting the target music, the audio of the multiple target instrument tracks is played simultaneously.
5. The method according to claim 3, characterized in that, The step of splicing the target music segment according to the timestamp to obtain the target music includes: The target music segments are sorted according to the timestamps to obtain sorted target music segments; Process overlapping or blank segments of the sorted target music segments, and then splice the sorted target music segments in sequence to obtain the target music.
6. The method according to claim 1, characterized in that, The method includes the construction of a music source separation model, which includes a depthwise separable convolutional model, a knowledge distillation model, and an extraction model; The step of separating the music source signal using a music source separation model to obtain multiple instrument tracks of the music source signal includes: The original depthwise separable convolutional model is used as the student model in the knowledge distillation model. The trained teacher model is then imitated using knowledge distillation techniques to obtain the depthwise separable convolutional model. The music source signal is extracted from the music source signal using the depthwise separable convolution model, and the music source signal is initially separated based on the music source signal to generate multiple original music source signals; The extraction model is used to optimize the multiple original sound source signals to obtain the instrument track corresponding to each music source signal.
7. A music synthesis device, characterized in that, The device includes: The acquisition module is used to acquire a music source signal, wherein the music source signal contains at least two instrument tracks; The separation module is used to separate the music source signal using a music source separation model to obtain multiple instrument tracks of the music source signal. The music source separation model is a lightweight model of an embedded system. The editing module is used to identify the target track frame of each instrument track, perform editing operations on the target track frame, and obtain the target instrument track. The processing module is used to merge multiple target instrument tracks and export the target music.
8. A portable electronic device, characterized in that, include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
10. A vehicle, characterized in that, include: The music synthesis apparatus as described in claim 7; Alternatively, the portable electronic device as described in claim 8; Alternatively, the computer-readable storage medium as described in claim 9.