Method for vem-token beat capture and alignment model construction

The VEM-Token beat capture and alignment model solves the problem of beat capture and alignment in existing technologies, achieving accurate beat synchronization and alignment of vocal files, supporting the development of music AI applications, and expanding the application scope of AI.

CN120748450BActive Publication Date: 2025-11-21GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511249168.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-11-21
Estimated Expiration
2045-09-03

AI Technical Summary

Technical Problem

In existing technologies, beat capture methods are relatively simple, and users cannot achieve synchronous beat capture and beat alignment when imitating sample files. There is a lack of systematic solutions, especially in multimodal vocal emotion processing, where it is difficult to achieve accurate beat capture and alignment.

Method used

The VEM-Token beat capture and alignment model is adopted. By setting the beat capture and beat alignment model, the vocal files are divided into VEM-Token sequences using the VEM-Token vocal emotion multimodal model. Combined with techniques such as multiple filter layers, harmonic impact sub-model, time-frequency convolutional network, and dynamic time warping sub-model, the fine-tuning alignment of the start and end points of the beats is achieved, ensuring the beat synchronization between user files and sample files.

Benefits of technology

It achieves accurate beat capture and alignment of vocal files, supports the development of music agents, and provides a model access method for music/vocal music to form music AI application agents, solving the problem of beat capture and alignment in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748450B_ABST
    Figure CN120748450B_ABST
Patent Text Reader

Abstract

The method of VEM-Token beat capture and alignment model construction is based on the VEM-Token vocal emotion multi-modal model method, which uses musical beats to further innovate the division of vocal files into VEM-Token units. The core of this method is to establish a beat model, a beat capture model, and a beat alignment model for vocal files. The former separates the sample vocal file into singing, accompaniment, and emotional fluctuations through multiple filters, captures the start and end points of the beat in the frequency spectrum format file, and the latter uses start and end point fine-tuning models to align the user's imitation file with the sample file. The beat is captured using models including beat base model, harmonic impact, joint learning, harmonic frequency layering, dynamic time warping, etc. The construction of models such as base model, start point fine-tuning and end point fine-tuning, full alignment verification, beat editor, free beat, repeat alignment, and communication interface protocol makes this method suitable for Agent music intelligent agents and AI music applications.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, specifically to the model construction and processing of AI intelligent agent Agent, AI music system and voice recitation, especially when people imitate sample vocal music or sample recitation, the processing of beat capture and beat alignment is carried out to realize and optimize the imitation effect. The method of VEM-Token beat capture and alignment model construction is based on the VEM-Token vocal emotion multi-modal model method, which uses music beats to divide vocal music files into VEM-Token word units. BACKGROUND

[0002] Based on the invention of the inventor's already granted Chinese invention patent "VEM-Token vocal emotion multi-modal tokenized song and accompaniment deep learning method, CN120126506" (hereinafter referred to as "VEM-Token vocal emotion multi-modal model"), in the field of artificial intelligence, it is different from the traditional NLP-Token (Natural-Language-Processing Token) word division method, and it is the first time to divide information word units according to music beats, called vocal emotion multi-modal token, simply referred to as VEM-Token (Vocal-Emotion-Multimodal Token).

[0003] In the CN120126506 invention patent, the inventor first spectrally processes vocal music files, detects beats, divides the spectrally processed vocal music files into VEM-Token sequences according to vocal music beats, establishes VEM coordinate system, VEM function and VEM library according to multi-modal such as lyrics, song, accompaniment, singer emotion, accompaniment emotion, video, image, etc., performs VEM-Token recognition, separates song stream and accompaniment stream, scores multi-modal emotions of vocal samples according to vocal experts, obtains VEM parameters by using supervised learning and deep learning algorithm, and learns multi-modal emotions of vocal samples. For other vocal works, it can identify vocal multi-modal emotions, output lyrics score, VEM-Token song score, VEM-Token accompaniment score and VEM-Token score. The present application as a patent pool patent of CN120126506 can be connected to AI system or independently developed application system, and developed into a vocal intelligent agent Agent that can listen to songs and recognize scores.

[0004] The VEM-Token concept mentioned in the present application is based on the basic concept and steps defined in the invention patent CN120126506, unless otherwise emphasized or specifically defined in the present application. Among them, the "vocal music file", "music file", "song", "music" mentioned in the present application file have the same meaning unless otherwise stated.

[0005] At present, the research and application of various large models are booming. For example, OpenAI, DeepSeek, Google Gemini, Kimi, big model, and Wenxin Yiyang, etc. After connecting the front end and back end, various intelligent agents are formed. The present application aims to access these large models and realize two-way communication with them to form music / vocal-based artificial intelligence applications, even music AI agents, to expand the application of AI and provide powerful innovation and support.

[0006] Disadvantages of prior art methods

[0007] Since the invention patent CN120126506 is the first to propose VEM-Token multi-modal tokenization processing of music information, there are still the following deficiencies in music tempo and its alignment:

[0008] 1. The tempo capture method is relatively simple. For sample files, how to make the user file complete the synchronous capture tempo when the user imitates the sample file, and no method for tempo capture is given.

[0009] 2. For sample files, how to make the user file complete the tempo alignment with the sample file, no definition and optimization is given.

[0010] 3. There is no system solution for the tempo capture of VEM-Token and the tempo alignment of double files. SUMMARY

[0011] In view of the deficiencies of the prior art, the present application proposes a completely new and complete innovative solution to realize the capture and recognition of vocal tempo and the alignment of music file 2 to music file 1. To solve the difficulties of conventional NLP-Token methods based on text for multi-modal vocal tempo capture and recognition and tempo alignment, to achieve the purpose and intention of the present application.

[0012] The purpose and intention of the present application is achieved by using the following technical solutions and working steps:

[0013] 1. Basic scheme implementation steps

[0014] The present application is a method for constructing a VEM-Token tempo capture and alignment model, which includes but is not limited to the following steps:

[0015] ST100: For vocal music files, set the beat model including beat capture and beat alignment according to the VEM-Token vocal emotion multi-modal model, capture the beat of the vocal music file, divide the vocal music file into VEM-Token sequences according to the beat, and mark the position of the start of the beat and the end of the beat in each VEM-Token.

[0016] ST200: Set the start point alignment model, including:

[0017] The vocal music files include but are not limited to sample files and user files produced by the user singing the sample files, which are divided into sequences and sequences, for each VEM-Token in the sequence, according to the start point of each , use the start point fine-tuning step to adjust the start point of at the corresponding position one by one, so as to align with the start point of at the corresponding position.

[0018] For each segment of the vocal music file including but not limited to the loop segment, take the first segment as a reference, starting from the second segment, use the start point fine-tuning step to adjust the start point of each VEM-Token of each segment one by one, so as to align with the start point of the VEM-Token at the corresponding position of the first segment, until all loop segments end.

[0019] ST300: Set the end point alignment model, including but not limited to:

[0020] According to the end point of each , use the end point fine-tuning step to adjust the end point of at the corresponding position one by one, so as to align with the end point of at the corresponding position.

[0021] For each segment of the loop segment, take the first segment as a reference, starting from the second segment, use the end point fine-tuning step to adjust the end point of each VEM-Token of each segment one by one, so as to align with the end point of the VEM-Token at the corresponding position of the first segment, until all loop segments end.

[0022] 2. Basic sub-model steps

[0023] On the basis of the foregoing basic scheme, the present application in establishing the beat sub-model specifically includes but is not limited to one or more combinations of the following steps or methods:

[0024] ST110: Adopt VEM-Token vocal emotion multi-modal model including but not limited to VEM processor, set multiple filter layers, convert sample files and user files into spectral format files to separate human singing, instrumental accompaniment and emotional overtones and fluctuations, including but not limited to:

[0025] Singing filter, filter frequency range 80-1500Hz, stopband attenuation ≥ 48dB / oct.

[0026] Accompaniment filter, filter frequency range 25-3000Hz, stopband attenuation ≥ 36dB / oct.

[0027] Emotional overtone filter, filter frequency range 1400-12000 Hz.

[0028] Emotional fluctuation filter, filter frequency range 0.01-1.0 Hz.

[0029] oct is the rule of octave in music theory, that is, the audio frequency difference of each oct is one time.

[0030] ST111: Establish beat model formula including but not limited to:

[0031] 2.1

[0032] 2.2

[0033] 2.3

[0034] 2.4

[0035] 2.5

[0036] 2.6

[0037] 2.7

[0038] 2.8

[0039] 2.9

[0040] 2.10

[0041] 2.11

[0042] 2.12

[0043] 2.13

[0044] in:

[0045] Formula 2.1 is the sample file Sequence, equivalent to sample file The sequence, where N is the total number of beats in the sample file;

[0046] Formula 2.2 is for user files Sequence, equivalent to user files The sequence, where N is the total number of beats in the user file.

[0047] Formulas 2.3 and 2.4 are respectively sequence, The content of the nth beat in the sequence, where n is the beat number from 1 to N.

[0048] Formulas 2.5 and 2.6 indicate that the end position of the nth beat is equal to the starting position of the (n+1)th beat.

[0049] , Sequences , The starting pointer and position of the nth beat.

[0050] , The sequences are respectively , The pointer and position of the end point of the nth beat are also the pointer and position of the start point of the (n-1)th beat.

[0051] Formulas 2.7 and 2.8 are respectively , The duration of a beat.

[0052] Formulas 2.9 and 2.10 are respectively sequence, The contents of the sequence.

[0053] Formula 2.11 represents the starting time difference between the sample file and the user file on the nth beat.

[0054] Formula 2.12 represents the time difference between the sample file and the user file at the end of the nth beat.

[0055] In formula 2.13 It is the minimum absolute value of all time-lapse errors, that is, the minimum value among the absolute values ​​of the logical OR of all start time differences and all end time differences.

[0056] ST112: Beats include, but are not limited to, rhythms. The unit of time is the number of beats or rhythms per minute, and the duration is measured in milliseconds (ms).

[0057] ST113: The rhythm includes uniform rhythm that keeps the same rhythm before and after in a period of time, and mutation rhythm that starts to mutate to another uniform rhythm after a uniform rhythm, and the error between a uniform rhythm is less than 5mS.

[0058] ST114: According to music theory, the beat includes but is not limited to strong, weak, and secondary strong beat types.

[0059] ST115: According to music theory, the vocal music file VEM-Token sequence is also divided into measures, each measure includes one or more beats, and the timing unit of the measure is determined according to the beat.

[0060] ST116: According to music theory, in the vocal music file, a cycle section includes but is not limited to one or more measures.

[0061] 3. The beat capture model includes a harmonic shock sub-model

[0062] On the basis of the foregoing scheme, in the aspect of the beat capture model, the present application includes but is not limited to the harmonic shock sub-model HPS, and specifically includes but is not limited to one or more combinations of the following steps or methods:

[0063] ST120: For a music file including but not limited to percussion music, the spectral format file is separated into a harmonic component spectrogram including but not limited to sustained sound and a shock component spectrogram including but not limited to transient sound.

[0064] ST121: According to the shock component spectrogram, a spectral energy mutation point is detected as a rhythm pointer candidate point.

[0065] ST122: A peak value threshold condition is set, and a peak value point meeting the condition is detected as a rhythm pointer candidate point.

[0066] ST123: According to music theory rules, the rhythm pointer candidate point is detected with reference to the harmonic component spectrogram, non-main rhythm candidate points are deleted, and main rhythm points are reserved.

[0067] ST124: With reference to the harmonic component spectrogram, the main rhythm points are aligned, and beat starting points are generated.

[0068] 4. The beat capture model includes a joint learning sub-model

[0069] On the basis of the foregoing scheme, in the aspect of the beat capture model, the present application includes but is not limited to the joint learning sub-model, and specifically includes but is not limited to one or more combinations of the following steps or methods:

[0070] ST130: Extract multiple spectral features from the input using a time-frequency convolutional network, model the time sequence through a bidirectional recurrent network, and output a beat position heat map and a time signature classification probability in parallel in the output layer. The heat map is weighted according to the time signature type, and the tempo pointer is extracted.

[0071] ST131: Fuse the three channels of Mel-spectrum, Constant-Q transform spectrum, and instantaneous phase spectrum to construct a multi-channel time-frequency input, input it into a convolutional recurrent neural network (CRNN), and the output layer includes a first branch that outputs the probability of the existence of a beat at each time frame, and a second branch that outputs the time signature classification probability.

[0072] ST132: Perform weight correction according to the predicted time signature type, including but not limited to: if it is 4 / 4 time, multiply the probability values of the 1st and 3rd beats by a weighting coefficient, and if it is 3 / 4 time, multiply the probability value of the 1st beat by a weighting coefficient.

[0073] ST133: Take the local maximum point of the extracted and corrected probability as the final tempo pointer.

[0074] 5. The beat capture model includes a harmonic frequency hierarchical sub-model

[0075] On the basis of the foregoing scheme, the beat capture model includes but is not limited to a harmonic frequency hierarchical sub-model, which specifically includes but is not limited to one or more combinations of the following steps or methods:

[0076] ST140: Use a convolutional neural network to convert sample files and user files into multiple hierarchical spectral format files, including but not limited to beat frequency, vocal fundamental frequency, harmonic fundamental frequency, vocal overtone, instrument fundamental frequency, and instrument overtone, each of which includes but is not limited to:

[0077] The beat frequency is 0.5-20 Hz.

[0078] The vocal fundamental frequency is 80-1400 Hz.

[0079] The harmonic fundamental frequency is 100-500 Hz.

[0080] The vocal overtone is 1.5-8 kHz.

[0081] The instrument fundamental frequency is 27.5-4186 Hz.

[0082] The instrument overtone is 4-16 kHz.

[0083] ST141: Use an adaptive convolutional network to calculate the tempo pointer based on the upper and lower limits of each hierarchical frequency.

[0084] 6. The beat capture model includes a dynamic time warping sub-model

[0085] On the basis of the foregoing scheme, the present application includes but is not limited to a beat capture model, specifically including but not limited to the following one or more combinations of steps or methods:

[0086] ST150: Extracting phonemes, beats and lyric sentences in the singing voice from the sample file and the user file, and positioning including but not limited to:

[0087] Phoneme boundary anchor points, locating initial consonant-vowel transition points or plosive starting points through a hidden Markov model (HMM) as rhythm pointers.

[0088] Beat strength anchor points, marking strong beats and secondary strong beats based on energy envelope maximum points as rhythm pointers.

[0089] Sentence gas port anchor points, detecting <-40 dB silent sections and spectral mutation points as rhythm pointers.

[0090] The beat time error is not more than 50 milliseconds.

[0091] 7. Beat alignment model

[0092] On the basis of the foregoing scheme, the present application includes but is not limited to a beat model, specifically including but not limited to the following one or more combinations of basic model steps or methods:

[0093] ST210: Based on the setting, for a beat , the alignment state of the user beat and the sample beat includes but is not limited to:

[0094] When , the start pointer state of the nth beat of the user file and the sample file is aligned.

[0095] When , the start pointer state of the nth beat of the user file and the sample file is lagging behind.

[0096] When , the start pointer state of the nth beat of the user file and the sample file is ahead.

[0097] The beats of the vocal music file include but are not limited to: marches, lyrical songs, strong rhythm songs, and diffuse song with unclear rhythm.

[0098] ST220: The value judgment range of beat error includes but is not limited to:

[0099] For singing voice, .

[0100] For accompaniment, .

[0101] Regarding emotions, .

[0102] For vocal files of lyrical songs with strong or slow tempos, the tempo error is reduced or increased according to the song, but the value is no more than 20% of the corresponding tempo.

[0103] For loose boards Extend or cancel.

[0104] 8. Steps for fine-tuning the starting point and the ending point

[0105] Based on the aforementioned scheme, the present invention also includes, but is not limited to, for a beat numbered n, sequentially performing a start-point fine-tuning step and an end-point fine-tuning step, specifically including, but not limited to, one or more combinations of the following steps or methods:

[0106] ST230: For beat start pointer states that are lagging, the start fine-tuning steps include, but are not limited to:

[0107] when Then shift to the left The translation time is ,

[0108] when Then in Add a length of on the left Ornaments, including but not limited to vibrato, glissando, breathiness, or, in Internal copy The duration of the content can be extended, or, in Internal cloning The duration of the content has been extended.

[0109] ST240: If the starting pointer state of the beat is ahead, the starting fine-tuning steps include, but are not limited to, shifting to the right. The translation time is , or, in Internal deletion The duration of the content has been extended.

[0110] ST310: Based on the completion of the initial fine-tuning steps Recalculate and calibrate the end pointer of the beat numbered n, and calculate the... The endpoint alignment state is obtained, and the endpoint lagging state and endpoint leading state are obtained.

[0111] ST320: For beat endpoint pointers that are lagging, endpoint fine-tuning steps include, but are not limited to:

[0112] when Then shift the endpoint to the left, and the shift time is... .

[0113] when Then compress The compression method is the same as the original. Compression at the same frequency.

[0114] ST330: For beat endpoint pointer states that are ahead, the endpoint fine-tuning steps include, but are not limited to, inserting a length of [unclear] after the endpoint. Ornaments, including but not limited to trills, glissandos, breath sounds, rests, or, in Internal copy The duration of the content can be extended, or, in Internal cloning The duration of the content has been extended.

[0115] 9. Steps for full beat alignment and verification

[0116] Based on the aforementioned scheme, and in the beat alignment model, the present invention further includes, but is not limited to, one or more combinations of the following steps or methods:

[0117] ST410: All Each beat in the sequence is a working set. Starting from the first beat, the start-point fine-tuning step and the end-point fine-tuning step are executed cyclically to complete the entire sequence. Sequence for Sequence alignment.

[0118] ST420: Based on Verification of all start and end points of the sequence The alignment of the start and end points of the sequence is determined by the user; any misaligned beats must be manually handled.

[0119] 10. Steps for using the beat editor

[0120] Based on the aforementioned solution, this invention also includes, but is not limited to, the steps or methods of a beat editor in relation to the beat alignment model:

[0121] ST430: The beat editor provides a user interface for all operations, including defining, editing, collecting sample files, recording user files, determining beats, and aligning beats.

[0122] ST440: The beat editor includes, but is not limited to, functions for automatically determining the beat, starting point calibration, ending point calibration, starting point alignment, ending point alignment, and beat data display. It also includes, but is not limited to, functions for manually adjusting and determining the beat, starting point calibration, ending point calibration, starting point alignment, ending point alignment, and beat data display based on the needs of vocal emotion.

[0123] ST450: change the tempo of the sequence according to the measure or cycle segment, make the beat modification and beat alignment across measures or across cycle segments.

[0124] ST460: according to the style of the vocal file, mark the equal beat in the VEM-Token sequence as uniform beat, mark the equal measure time interval in the VEM-Token sequence of the vocal file of the emotional style as local uneven beat, and mark the unequal measure time interval in the VEM-Token sequence of the vocal file of the allegro style as global uneven beat.

[0125] 11. Free beat step

[0126] On the basis of the foregoing scheme, the present application further includes but is not limited to the steps or methods of free beat:

[0127] ST500: according to the special processing of the user for the sound emotion, in the special processing of the local beat, the starting point of the beat is not aligned with the corresponding sequence, the ending point of the beat is not aligned with the corresponding sequence, and the user emotion is personalized and freely played in the special processing of the local beat, with the starting point and the ending point in advance or lag.

[0128] ST510: after the special processing of the local beat, the alignment of the starting point and the ending point of the beat is restored.

[0129] ST520: the beat editor supports editing and marking of the freely played beat.

[0130] 12. Repeat alignment step

[0131] On the basis of the foregoing scheme, the present application further includes but is not limited to the steps or methods of repeat alignment step:

[0132] ST600: the beat and alignment model further includes the step of repeat alignment.

[0133] ST610: when the beat alignment model is submitted to the subsequent large model once, the result is traced back to the beat alignment model after the large model runs the processing result, to recheck the beat alignment scheme, and if there is a deviation in the beat alignment, the operation of re-beat alignment is performed again.

[0134] ST620: the step of providing the beat alignment synchronization signal to the subsequent large model application, the synchronization signal includes but is not limited to beat number, beat type, and beat width information.

[0135] ​​ST630: According to the complete music file, the beat-aligned music file including the five-line staff, the tablature and the beat score is outputted.

[0136] 13. Communication interface protocol steps

[0137] On the basis of the foregoing scheme, the present application further includes but is not limited to the steps or methods of the communication interface protocol:

[0138] ST700: According to the beat-aligned synchronization signal provided to the subsequent large model application, through the bidirectional communication protocol and bidirectional interface protocol for the intelligent agent, the AI intelligent agent Agent is formed together with the subsequent large model.

[0139] ST800: According to the beat-aligned synchronization signal provided to the subsequent large model application, through the bidirectional communication protocol and bidirectional interface protocol for the independent AI music system, the AI music system is formed.

[0140] ST900: The bidirectional communication interface protocol includes: the uplink data, the downlink data and the handshake rule of the present method for the intelligent agent Agent or the AI music system.

[0141] 14. Purpose and intention of the application

[0142] The purpose and intention of the VEM-Token beat capture and alignment model construction method of the present application are:

[0143] The beat capture model of the music file is created to accurately capture the VEM-Token beat of the music file.

[0144] The beat alignment model of the music file is created to make the beat of the music file 2 (user imitation sample file ) aligned with the beat of the music file 1 (sample file ).

[0145] The beat editing model is created to facilitate the user to make personalized beat alignment and adjustment.

[0146] An innovative music / vocal music model is provided, which accesses a large model to form a music AI application intelligent agent Agent or a dedicated application system, thereby providing strong support for expanding the application of AI.

[0147] 15. Beneficial effects of the application

[0148] 1. The purpose and intention of the application are achieved.

[0149] 2. The problem of the existing artificial intelligence being unable to recognize the vocal music beat capture and beat alignment is solved.

[0150] 3. The music intelligent agent Agent is supported.

[0151] 4. Innovate a music / vocal music oriented model, access a large model to form a music AI application agent, or a dedicated application system, to provide strong support for expanding the application of AI. BRIEF DESCRIPTION OF DRAWINGS

[0152] LIST OF DRAWINGS

[0153] Figure 1 : Music beat division and alignment schematic diagram

[0154] Figure 2 : VEM processor schematic diagram

[0155] Figure 3 : Harmonic impact sub-model HPS flow reference schematic diagram

[0156] Figure 4 : High harmonic schematic diagram

[0157] Figure 5 : VEM-Token segmentation schematic diagram

[0158] Figure 6 : Joint learning sub-model flow reference schematic diagram

[0159] Figure 7 : Harmonic frequency hierarchical sub-model flowchart

[0160] Figure 8 : Dynamic time warping sub-model DTW flowchart

[0161] Figure 9 : VEM-Token comprehensive beat score diagram.

[0162] DETAILED DESCRIPTION OF DRAWINGS

[0163] See the specific embodiments. DETAILED DESCRIPTION

[0164] The present application is a patent pool patent of the already authorized Chinese invention patent "VEM-Token vocal emotion multi-modal tokenization song and accompaniment deep learning method, CN120126506", focusing on the music beat capture invention idea, and the user file beat alignment invention idea facing sample files, and basic innovation is made.

[0165] The purpose and intention of the present application can be realized by the following specific embodiments. It should be particularly noted that since the specific embodiments have specific purposes and industrial applicability, the embodiments cannot include all the features and steps of the present application, nor are they a limitation of the present application. The description of the claims of the present application is the summary of the invention.

[0166] This example is one of the examples of the present invention.

[0167] The detailed description of the present invention is as follows:

[0168] An AI system and AI agent for precise implementation of music beat capture and beat alignment

[0169] Innovative VEM-Token beat capture and alignment model construction method.

[0170] Diagram explanation

[0171] The present embodiment mainly includes but is not limited to the following main schematic drawings, which are: Figures 1 to 9 .

[0172] Implementation step explanation

[0173] The method steps of the present embodiment mainly include steps 1 to 13. Unless otherwise specified, the step numbers of these 13 parts do not have a sequence, and the combination of these 13 parts is not required for any embodiment. In addition, each of the 13 parts includes several sub-steps. Unless otherwise specified, these sub-steps are not completely required, and their sequence is not necessary, but according to the needs of some specific tasks, the patent implementer makes optimized and further choices.

[0174] The specific work steps are as follows:

[0175] 1. VEM-Token beat alignment model basic scheme implementation steps

[0176] The present invention as a VEM-Token beat capture and alignment model construction method includes at least but not limited to the following steps:

[0177] ST100: For vocal files, set a beat model including beat capture and beat alignment according to the VEM-Token vocal emotion multi-modal model, capture the beats of the vocal files, divide the vocal files into VEM-Token sequences according to the beats, and mark the positions of the start of the beat and the end of the beat in each VEM-Token.

[0178] ST200: Set a start alignment model, including:

[0179] Divide the vocal files including but not limited to sample files and user imitation sample file singing generated user files into sequences and sequences, for each VEM-Token in the sequence, according to the start of each , use the start fine-tuning step to adjust the corresponding position of the first segment, so that the start point of the VEM-Token at the corresponding position is aligned with the start point of the VEM-Token at the corresponding position of the first segment. of the first segment.

[0180] For each segment of the loop segment, with the first segment as a reference, starting from the second segment, the start point of each VEM-Token of each segment is adjusted one by one by using the start point fine-tuning step, so that the start point of the VEM-Token at the corresponding position is aligned with the start point of the VEM-Token at the corresponding position of the first segment, until the end of all loop segments.

[0181] ST300: Set the end point alignment model, including but not limited to:

[0182] According to the end point of each , the end point of the corresponding position is adjusted one by one by using the end point fine-tuning step, so that the end point of the VEM-Token at the corresponding position is aligned with the end point of the VEM-Token at the corresponding position of the first segment.

[0183] For each segment of the loop segment, with the first segment as a reference, starting from the second segment, the end point of each VEM-Token of each segment is adjusted one by one by using the end point fine-tuning step, so that the end point of the VEM-Token at the corresponding position is aligned with the end point of the VEM-Token at the corresponding position of the first segment, until the end of all loop segments.

[0184] Among them, the model of VEM-Token vocal emotion multi-modal refers to the model in "VEM-Token vocal emotion multi-modal tokenization song and accompaniment deep learning method, CN120126506", which specifically includes 1-4:

[0185] 1. Record emotions using one or more modalities, mark the vocal emotion multi-modal as VEM, and construct VEM classification, VEM coordinate system, VEM function and VEM library. Vocal emotion includes one or a combination of joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, emotion, and enmity. Multi-modal includes one or a combination of lyrics, song, accompaniment, vocal style, music, emotion basis, accompaniment instrument, video, and image. VEM coordinate system includes a coordinate axis system established according to independent emotions, opposite emotion pairs, and related opposite emotion groups.

[0186] 2. Collect vocal samples according to VEM classification, and perform emotion evaluation on the vocal samples by human vocal experts in terms of song and accompaniment. Use supervised learning and deep learning to train VEM function to obtain VEM parameters and add them to VEM library.

[0187] 3. Use VEM processor to beat mark the vocal file, and separate the song stream and the accompaniment stream. According to the beat, the vocal file is segmented into VEM-Token, the song stream is converted into sequence, and the accompaniment stream is converted into​​ The sequence is then added to the preprocessing library.

[0188] 4. Using deep learning, lyrics scores, VEM-Token vocal scores, VEM-Token accompaniment scores, and VEM-Token sheet music are generated respectively.

[0189] In this invention application, the beat model includes a beat capture model and a beat alignment model, and the beat alignment model includes a start-point alignment model and an end-point alignment model.

[0190] Figure 1 Diagram showing the division and alignment of musical beats.

[0191] exist Figure 1 Including the lexical units of the sample file ( ) and user file lexicals ( The user file is a vocal recording of the user imitating the sample file. According to music theory, the user file should be as consistent as possible with the sample file. In this invention, this consistency primarily means that the beats must be identical; that is, the start and end points of the beats in the user file must be aligned with the start and end points of the beats in the sample file. It should be noted that, based on the working principle of Von der Von's serial array, the word tokens parsed from the vocal file exist in the form of a numerical sequence (such as an array), and are divided into VEM-Token sequences according to the musical beat format. The sample file is divided into... Sequences, user files are divided into A sequence. Each lexical unit defines a start and an end point. The start point of the next lexical unit coincides with the end point of the previous lexical unit. A lexical unit sequence is defined as follows: Figure 1 The subscript shown.

[0192] This step includes a beat start alignment model and a beat end alignment model. The beat start and beat end alignment models actually involve beat start and beat end capture processes. This is because music files may contain "rests" or "breaths," which are represented in the array signal (spectral format file) as array elements with a semaphore value of 0 or close to 0. Therefore, to avoid such interference, this invention employs a start-point fine-tuning step and an end-point fine-tuning step to further accurately capture the beat start and beat end.

[0193] Additionally, a music file may contain "loop sections." Within a loop section, the score and rhythm of the next section are aligned with the previous section, including both identical and slightly different parts. For identical parts, the start and end points of the beats in the next section need to be aligned with the previous section. For different parts, the start and end points should be followed according to the actual start and end points of the sample file.

[0194] It should be noted that, generally in the sample file, since it is assumed that the sample file is a relatively professional singer singing and recording strictly according to the rules of music theory, the beat is accurate. However, the user may not be a professional singer, or a less professional music lover, so the starting point of the beat and the ending point of the beat cannot be strictly sung according to the rules of music theory, that is, the starting point of the beat and the ending point of the beat are not strictly consistent with the sample file. Therefore, in this case, the capture and alignment of the starting point of the beat and the ending point of the beat of the user file need to be strictly checked and modified.

[0195] 2. Implementation steps of the basic sub-model of the VEM-Token beat model

[0196] On the basis of the foregoing scheme, the present application in the aspect of beat sub-model specifically includes but is not limited to one or more combinations of the following steps or methods:

[0197] ST110: Using the VEM-Token vocal music emotion multi-modal model including but not limited to the VEM processor, setting multiple filter layers, converting the sample file and the user file into a spectrum format file to separate the human singing voice, the accompaniment sound of the instrument, and the emotion overtone and fluctuation, specifically including but not limited to:

[0198] The singing voice filter, the filter frequency range is 80-1500Hz, and the stopband attenuation is ≥48dB / oct.

[0199] The accompaniment filter, the filter frequency range is 25-3000Hz, and the stopband attenuation is ≥36dB / oct.

[0200] The emotion overtone filter, the filter frequency range is 1400-12000 Hz, the 3kHz gain is +6dB, and the 100Hz attenuation is ≤-36dB.

[0201] The emotion fluctuation filter, the filter frequency range is 0.01-1.0 Hz.

[0202] oct is the octave rule in music theory, that is, the audio frequency of each oct differs by a factor of two.

[0203] ST111: Establishing a beat model formula including but not limited to:

[0204] 2.1

[0205] 2.2

[0206] 2.3

[0207] 2.4

[0208] 2.5

[0209] 2.6

[0210] 2.7

[0211] 2.8

[0212] 2.9

[0213] 2.10

[0214] 2.11

[0215] 2.12

[0216] 2.13

[0217] Where:

[0218] Equation 2.1 is the n-th beat content in the sample file's sequence, equivalent to the n-th beat content in the sample file's sequence, N is the total number of beats in the sample file.

[0219] Equation 2.2 is the n-th beat content in the user file's sequence, equivalent to the n-th beat content in the user file's sequence, N is the total number of beats in the user file.

[0220] Equations 2.3, 2.4 are the n-th beat start pointer and position in the sequence, respectively, n is the beat sequence number, from 1 to N.

[0221] Equations 2.5, 2.6 are the n-th beat end position, which is equal to the start position of the n+1-th beat.

[0222] , are the n-th beat start pointer and position in the , sequence, respectively.

[0223] , are the n-th beat end pointer and position in the , sequence, respectively, and also the start pointer and position of the n-1-th beat.

[0224] Equations 2.7, 2.8 are the n-th beat start pointer and position in the 、 The time length of a beat.

[0225] Equations 2.9 and 2.10 are respectively The sequence, The content of the sequence.

[0226] Equation 2.11 is the start time difference of the sample file and the user file at the nth beat.

[0227] Equation 2.12 is the end time difference of the sample file and the user file at the nth beat.

[0228] In equation 2.13, is the minimum absolute value of all beat errors, that is, the minimum value in the absolute value of the logical OR of all start time differences and all end time differences.

[0229] ST112: A beat includes, but is not limited to, a rhythm, and the time unit is the number of beats or rhythms per minute, or the number of milliseconds.

[0230] ST113: A rhythm includes a uniform rhythm that keeps the same rhythm before and after in a period of time, and a mutation rhythm that starts to mutate into another uniform rhythm after a uniform rhythm, and the error between a uniform rhythm is less than 5 milliseconds.

[0231] ST114: According to music theory, a beat includes, but is not limited to, strong, weak, and secondary strong beat types.

[0232] ST115: According to music theory, the vocal file VEM-Token sequence is also divided into measures, each measure includes more than one beat, and the time unit of a measure is determined according to a beat.

[0233] ST116: According to music theory, in a vocal file, it includes, but is not limited to, a loop section, and a loop section includes, but is not limited to, more than one measure.

[0234] It should be noted that only four filters are designed here, but this does not mean a limitation of the present application. Users of the present application can add or reduce them according to their own applications. For example, for Italian tenor bel canto songs, for children's songs, bass songs, and for Chinese folk songs, etc., the corresponding filters can be set up. The bandpass range of the filter can also be adjusted according to the specific frequency range.

[0235] Regarding formulas 2.1 to 2.13, they are only one of the methods for establishing the mathematical model of the present application and are not a limitation of the present application. Regarding the beat, in music theory, there are also some extended concepts and definitions, and the user of the present application should understand that the present application is around the cross-disciplinary innovative embodiments in music and vocal music, and is not a complete list and traversal of music and vocal music inventions. The user should make corresponding applications and modifications according to this idea.

[0236] Figure 2 : VEM processor schematic diagram

[0237] In the figure, from left to right, the sample file and the user file are input into the left frame and the right frame in sequence. In the left frame, various filters and beat capture model sets are mainly included. In the right frame, various beat capture model sets and beat alignment model sets are mainly included. The finally generated VEM-Token sequence is provided with accurate beat marks, and the data stream signal is connected to a large model or an AI music application.

[0238] 3. The beat capture model includes a harmonic percussive sub-model HPS

[0239] On the basis of the foregoing scheme, the present application includes but is not limited to the harmonic percussive sub-model HPS in the beat capture model, and specifically includes but is not limited to one or a combination of the following steps or methods:

[0240] ST120: For a music file including but not limited to percussion music, the spectral format file is separated into a harmonic component spectrogram including but not limited to sustained sound and a percussive component spectrogram including but not limited to transient sound.

[0241] ST121: According to the percussive component spectrogram, a frequency spectrum energy mutation point is detected as a rhythm pointer candidate point.

[0242] ST122: A peak value threshold condition is set, and a peak value point meeting the condition is detected as a rhythm pointer candidate point.

[0243] ST123: According to the rules of music theory, the harmonic component spectrogram is referred to, the rhythm pointer candidate points are detected, non-main rhythm candidate points are deleted, and main rhythm points are reserved.

[0244] ST124: Referring to the harmonic component spectrogram, the main rhythm points are aligned, and a beat starting point is generated.

[0245] It should be noted that the harmonic percussive sub-model HPS (Harmonic Percussive Separation, HPS) has the basic principle of decomposing sound waves from harmonic components and percussive wave components in two directions, and the target of the HPS algorithm is to separate them by mathematical means by using the different structural characteristics of the two components in the time-frequency domain.

[0246] The application of HPS is not limited to music files including percussion, and some music does not include percussion, but has obvious rhythm, such as Mongolian short tunes, horseback songs, children's songs, marches, etc., which are also suitable for using the HPS sub-model to capture the beat, form the rhythm pointer, and form the beat starting point.

[0247] Figure 3 Harmonic impact sub-model HPS flow reference schematic diagram

[0248] This is a reference signal flow diagram of HPS. In the diagram, STFT is a short-time Fourier transform (STFT, short-term Fourier transform), which is a mathematical transform related to the Fourier transform, used to determine the frequency and phase of the local region of the time-varying signal. Finally, the VEM-Token sequence is output.

[0249] Figure 4 Harmonic impact sub-model HPS flow reference schematic diagram

[0250] In the diagram, the red sinusoidal wave is the fundamental wave, the blue sinusoidal wave is the third harmonic wave, and the green sinusoidal wave is the fifth harmonic wave. In music, sound waves are not single-frequency sinusoidal waves, but contain rich high-order harmonics. In this diagram, only three frequency sinusoidal waves including the fundamental wave, the third harmonic wave and the fifth harmonic wave are drawn, and in the actual environment, a sound also includes a large number of other high-order harmonics.

[0251] Figure 5 VEM-Token segmentation schematic diagram

[0252] This is a spectrum diagram of an actual music recording. According to the HPS harmonic impact sub-model, the amplitude of the waveform in the diagram is used to detect the mutation point of the spectrum energy, and the beat starting point and the beat ending point are captured, as shown in the diagram.

[0253] 4. The beat capture model includes a joint learning sub-model

[0254] On the basis of the foregoing scheme, the present application includes but is not limited to a joint learning sub-model on the beat capture model, and specifically includes but is not limited to one or more combinations of the following steps or methods:

[0255] ST130: Extract multiple spectrum features using a time-frequency convolution network, model the time sequence through a bidirectional recurrent network, generate a beat position heat map and a time signature classification probability in parallel in the output layer, weight the heat map according to the time signature type, and extract the rhythm pointer.

[0256] ST131: three channels of fusion of mel-spectrogram, constant Q transform spectrogram and instantaneous phase spectrogram, construct a multi-channel time-frequency input, input to a convolutional recurrent neural network CRNN, the output layer includes a first branch in parallel, outputting the probability of the existence of the beat of each time frame, and a second branch in parallel, outputting the probability of the time signature classification.

[0257] ST132: weight correction according to the predicted time signature type, including but not limited to: if it is 4 / 4, the first and third beat probability values are multiplied by a weighting coefficient, and if it is 3 / 4, the first beat probability value is multiplied by a weighting coefficient.

[0258] ST133: taking the local maximum value point of the extracted corrected probability as the final rhythm pointer.

[0259] It should be noted that the joint learning sub-model is a multi-task learning prediction model. Its core innovation lies in using a shared backbone network to simultaneously complete the beat position regression and time signature classification tasks, and fusing the information of the two tasks through music theory rules to enhance each other.

[0260] Figure 6 : Joint learning sub-model process reference schematic diagram

[0261] This is one of the design examples of the joint learning sub-model process diagram of the end-to-end beat and time signature joint prediction model architecture, and the user of the present application should understand the basic program design, and design the own process diagram according to this step.

[0262] 5, the beat capture model includes a harmonic frequency hierarchical sub-model

[0263] On the basis of the foregoing scheme, the beat capture model of the present application includes but is not limited to a harmonic frequency hierarchical sub-model, specifically including but not limited to one or more combinations of the following steps or methods:

[0264] ST140: using a convolutional neural network to convert sample files and user files into a plurality of hierarchical spectrum format files, the harmonic frequencies including but not limited to beat frequencies, vocal fundamental frequencies, harmonic fundamental frequencies, vocal overtone frequencies, instrument fundamental frequencies and instrument overtone frequencies, respectively including but not limited to:

[0265] The beat frequency is 0.5-20 Hz.

[0266] The vocal fundamental frequency is 80-1400 Hz.

[0267] The harmonic fundamental frequency is 100-500 Hz.

[0268] The vocal overtone frequency is 1.5-8 kHz.

[0269] The instrument fundamental frequency is 27.5-4186 Hz.

[0270] The overtones of the musical instruments range from 4 to 16 kHz.

[0271] ST141: Employs an adaptive convolutional network to calculate the rhythm pointer based on the upper and lower limits of each layer's frequency.

[0272] The six frequency bands listed here are merely an example and are not the only or absolute setting. Users of this invention can freely adjust the number of frequency bands and the base frequency of each band according to their actual application and environment.

[0273] Figure 7 Flowchart of the steps to implement the layered sub-model of harmonic frequencies

[0274] The diagram shows how a CNN convolutional neural network is executed based on six frequency bands: beat frequency, fundamental voice frequency, harmonic fundamental frequency, voice overtones, instrument fundamental frequency, and instrument overtones. It's important to note that the size of the convolutional kernel is calculated by multiplying the frequency difference of each band by a parameter, such as 0.2. Users of this invention can adjust the kernel size to achieve optimal results for their specific applications.

[0275] 6. The beat capture model includes a dynamic time warping sub-model.

[0276] Based on the aforementioned scheme, the present invention includes, but is not limited to, the Dynamic Time Warping (DTW) sub-model in the beat capture model, specifically including, but not limited to, one or more of the following steps or methods in combination:

[0277] ST150: Extracts phonemes, beats, and lyric phrases from sample files and user files, locating, but not limited to:

[0278] Phoneme boundary anchor points are located using Hidden Markov Models (HMMs) to pinpoint the initial consonant-vowel transition point or the starting point of a plosive, thus serving as rhythm pointers.

[0279] The beat intensity anchor point, based on the maximum point of the energy envelope, marks the strong beat and the second strongest beat, and uses this as a rhythm indicator.

[0280] Sentence breath anchor points are used to detect silent segments <-40dB and abrupt changes in spectral entropy, which are then used as rhythm indicators.

[0281] The beat time error does not exceed 50 milliseconds.

[0282] Figure 8 A flowchart illustrating the implementation steps of the Dynamic Time Warping (DTW) submodel.

[0283] In the figure, the song "Embroidered Red Flag" is taken as a sample file, and the user file generated by imitating this song is inputted by using the two music files into the flow of the dynamic time warping (DTW) sub-model. It should be noted that in the present flow, the two innovations of "three-level anchor joint optimization" and "streaming phase continuity" are adopted.

[0284] 7. Beat alignment model

[0285] On the basis of the aforementioned beat capture scheme, according to the combination of one or more beat capture methods, the beat alignment model of the present application includes but is not limited to beat alignment, and specifically includes but is not limited to one or more combinations of the following basic model steps or methods:

[0286] ST210: Based on the setting, for a beat The alignment state of the user beat and the sample beat includes but is not limited to:

[0287] When The start pointer state of the nth beat of the user file and the sample file is aligned.

[0288] When The start pointer state of the nth beat of the user file and the sample file is lagging behind.

[0289] When The start pointer state of the nth beat of the user file and the sample file is ahead.

[0290] The beats of the vocal file include but are not limited to: march, lyrical song, strong rhythm song, and loose-plate song with no obvious rhythm.

[0291] ST220: Beat error The value judgment range includes but is not limited to:

[0292] For singing, .

[0293] For accompaniment, .

[0294] For emotion, .

[0295] For the vocal file of a strong rhythm or slow rhythm lyrical song, the beat error is reduced and enlarged according to the song, but the value is not greater than 20% of the corresponding beat.

[0296] For loose plate, extended or canceled.

[0297] It is important to emphasize here that beat alignment refers to:

[0298] 1. User files Sequence to sample file Align the sequences separately.

[0299] 2. Adjust only user files The start and end points of the sequence are usually not adjusted in the sample file.

[0300] 3. The range of values ​​for the beat error Δt can be adjusted based on the music beat attributes of the sample file. Under the special music attributes of user files, the adjustment range can exceed 20%.

[0301] 4. Under special emotional attributes, the adjustment of emotional rhythm attributes can be unrestricted by rhythm errors in order to ensure the musical effect.

[0302] 8. Steps for fine-tuning the starting point and the ending point

[0303] Based on the aforementioned scheme, the present invention also includes, but is not limited to, for a beat numbered n, sequentially performing a start-point fine-tuning step and an end-point fine-tuning step, specifically including, but not limited to, one or more combinations of the following steps or methods:

[0304] ST230: For beat start pointer states that are lagging, the start fine-tuning steps include, but are not limited to:

[0305] when Then shift to the left The translation time is ,

[0306] when Then in Add a length of on the left Ornaments, including but not limited to vibrato, glissando, breathiness, or, in Internal copy The duration of the content can be extended, or, in Internal cloning The duration of the content has been extended.

[0307] ST240: If the starting pointer state of the beat is ahead, the starting fine-tuning steps include, but are not limited to, shifting to the right. The translation time is , or, in Internal deletion The duration of the content has been extended.

[0308] ST310: Based on the completion of the initial fine-tuning steps Recalculate and calibrate the end pointer of the beat numbered n, and calculate the... the end-point alignment state, the end-point alignment state, the end-point lag state, the end-point lead state.

[0309] ST320: for the end-point pointer state of the beat lags, the end-point fine-tuning step includes but is not limited to:

[0310] When , the end-point is translated to the left, and the translation time is .

[0311] When , the is compressed, and the compression mode is to compress with the same frequency as the original .

[0312] ST330: for the end-point pointer state of the beat leads, the end-point fine-tuning step includes but is not limited to inserting a decorative sound with a length of after the end-point, including but not limited to tremolo, glissando, breath sound, rest sound, or copying the time length content of inside to extend, or cloning the time length content of inside to extend. Here, the user of the present application needs to pay special attention to:

[0313] 1. Fine-tune the start point first, and then fine-tune the end point. The order must not be reversed.

[0314] 2. The selection of the beat error Δt is determined by the style, property of the music and the user's requirement for beat alignment. The above is only the maximum error, and in general cases, it should be kept within the time resolution of the human ear, for example, less than

[0315] . 3. In any case where the start point fine-tuning of the VEM-Token occurs, the end point fine-tuning must be immediately executed, otherwise the beat may lose accuracy.

[0316] 9. Full-beat alignment and verification steps

[0317] On the basis of the foregoing scheme, the present application further includes but is not limited to one or more combinations of the following steps or methods:

[0318] ST410: take each beat in the entire sequence as a working set, and start from the first beat to execute the start point fine-tuning step and the end point fine-tuning step one by one in a loop to complete the alignment of the entire sequence to the

[0319] sequence.

[0320] ​​​ST420: According to the beginning and end of the sequence, check the alignment state of the beginning and end of the sequence, and manually handle the beat part that is not aligned by the user.

[0321] Here, the user of the present application needs to be reminded:

[0322] 1、This step can be used as an independent execution module, which needs to be executed in the case of possible influence on user file beat alignment, in order to check the beat alignment.

[0323] 2、Especially in the case of accessing late large models and AI intelligent Agent, this step needs to be executed to avoid the loss of accuracy and synchronization of beats due to the delay, illusion and divergence of large models.

[0324] 10、Beat editor step

[0325] On the basis of the foregoing scheme, the present application further includes but is not limited to one or more combinations of the following steps or methods:

[0326] ST430: The beat editor provides a man-machine interface for the user to define, edit, collect sample files, record user files, determine beats and align beats.

[0327] ST440: The beat editor includes but is not limited to the functions of automatically determining beats, beginning point calibration, end point calibration, beginning point alignment, end point alignment and beat data display, and also includes but is not limited to the functions of manually adjusting the determination of beats, beginning point calibration, end point calibration, beginning point alignment, end point alignment and beat data display based on the need for sound emotion.

[0328] ST450: According to the change of the rhythm of the sequence, the beat modification and beat alignment across the measure or across the cycle segment are carried out.

[0329] ST460: According to the style of the vocal music file, mark the beat time interval in the VEM-Token sequence as uniform beat, mark the beat time interval in the VEM-Token sequence of the vocal music file of the emotion type style as local uneven beat, and mark the beat time interval in the VEM-Token sequence of the vocal music file of the allegro type style as global uneven beat.

[0330] The beat editor includes:

[0331] 1、Based on a user interface design of a computer terminal and a mobile phone terminal, a drag cursor, a 3D dynamic interface and a voice control interface are used for control.

[0332] 2. The selection and editing of the communication protocol of the subsequent AI intelligent agent Agent to provide a complete communication interface.

[0333] 3. The communication interface of the subsequent independent AI music application is provided.

[0334] 11. Free rhythm step

[0335] On the basis of the foregoing scheme, the present application further includes but is not limited to one or more combinations of the following steps or methods:

[0336] ST500: According to the special processing of the user for the sound emotion, the special processing local rhythm in the sequence is not aligned with the starting point of the corresponding rhythm in the sequence, and the ending point of the rhythm is not aligned with the ending point of the corresponding rhythm in the sequence. The special processing local rhythm in the sequence is not aligned with the starting point of the corresponding rhythm in the sequence, and the ending point of the rhythm is not aligned with the ending point of the corresponding rhythm in the sequence. The special processing local rhythm in the sequence is not aligned with the starting point of the corresponding rhythm in the sequence, and the ending point of the rhythm is not aligned with the ending point of the corresponding rhythm in the sequence. The special processing local rhythm in the sequence is not aligned with the starting point of the corresponding rhythm in the sequence, and the ending point of the rhythm is not aligned with the ending point of the corresponding rhythm in the sequence.

[0337] ST510: After the special processing of the local rhythm, the alignment of the starting point and the ending point of the rhythm is restored.

[0338] ST520: The rhythm editor supports editing and marking of the free rhythm.

[0339] The user of the present application should understand that the free rhythm is very important when the singer performs personalized singing, especially in the case of personalized emotional expression, concert performance, and concert-like performance. In this case, the rhythm alignment of the present application should not be a constraint. For example, at the beginning, end and emotional climax of a song, the singer often needs to weaken, pause or even cancel the rhythm alignment when crossing rhythms, rhythms or even measures. For example, in a long Mongolian tune, the rhythm should support the free performance of the singer, and at the end of a lyrical song, the singer's free performance should also be supported.

[0340] 12. Repeat alignment step

[0341] On the basis of the foregoing scheme, the present application further includes but is not limited to one or more combinations of the following steps or methods:

[0342] ST600: The rhythm and alignment model further includes a repeat alignment step.

[0343] ST610: When the rhythm alignment model is submitted to the subsequent large model once, the result is traced back to the current rhythm alignment model after the large model runs and processes the result, to recheck the rhythm alignment scheme. If there is a deviation in the rhythm alignment, re-rhythm alignment operation is performed.

[0344] ST620: The step of providing a beat alignment synchronization signal to a subsequent large model application, including but not limited to beat number, beat type, and beat width information.

[0345] ST630: Outputting the beat-aligned score including staff, tablature, and beat score according to the complete music file.

[0346] When accessing subsequent complex applications, for example, some large model applications that are not very sensitive to time, due to the insertion of some unknown other applications, time movement occurs, so it is necessary to repeat the alignment of the accompaniment and to check the alignment of the beat. At the same time, it is suggested that the user of the present application needs to insert beat repeat calibration alignment in some possible positions in the overall application.

[0347] In addition, the user of the present application should also provide the beat number, beat type, and beat width information as part of the communication protocol to the backend application to achieve bidirectional efficient communication with the subsequent application.

[0348] Figure 9 : VEM-Token comprehensive beat score chart

[0349] In the figure, it includes staff, tablature, lyrics score, and beat score. It should be noted that this is just an example of a beat score chart, and the user of the present application can completely design other forms of beat score.

[0350] 13. VEM-Agent communication interface protocol steps

[0351] On the basis of the foregoing scheme, the present application further includes but is not limited to one or more combined steps or methods:

[0352] ST700: According to the beat alignment synchronization signal provided to the subsequent large model application, through the bidirectional communication protocol and bidirectional interface protocol for intelligent agents, an AI intelligent agent Agent is formed together with the subsequent large model.

[0353] ST800: According to the beat alignment synchronization signal provided to the subsequent large model application, through the bidirectional communication protocol and bidirectional interface protocol for independent AI music systems, an AI music system is formed.

[0354] ST900: The bidirectional communication interface protocol includes: the uplink data, downlink data, and handshake rules of the present method for intelligent Agent or AI music system.

[0355] The VEM-Agent communication interface protocol is further VEM-tokenized in reference to the Chinese invention patent "VEM-Token Vocal Emotion Multi-modal Tokenization Singing and Accompaniment Deep Learning Method, CN120126506", combined with the communication interface protocol of the system in the actual design, which can be used as a basis by users.

Claims

1. A method of VEM-token beat capture and alignment model construction, characterized by, Comprise: ST100: For vocal files, set the beat model including beat capture and beat alignment according to the VEM-Token vocal emotion multi-modal model, capture the beat of the vocal file, divide the vocal file into VEM-Token sequences according to the beat, and mark the position of the start of the beat and the end of the beat in each VEM-Token; ST200: Set the start point alignment model, comprising: Divide the sample file included in the vocal file and the user file sung by the user imitating the sample file into VEM-Token1 sequence and VEM-Token2 sequence respectively, and adjust the start point of VEM-Token2 at the corresponding position one by one according to the start point of each VEM-Token1, so as to align with the start point of VEM-Token1 at the corresponding position; For each segment of the loop segment included in the vocal file, take the first segment as a reference, and start from the second segment, adjust the start point of each VEM-Token of each segment one by one using the start point fine-tuning step, so as to align with the start point of VEM-Token at the corresponding position of the first segment, until all loop segments end; ST300: Set the end point alignment model, comprising: Adjust the end point of VEM-Token2 at the corresponding position one by one according to the end point of each VEM-Token1, so as to align with the end point of VEM-Token1 at the corresponding position; For each segment of the loop segment, take the first segment as a reference, and start from the second segment, adjust the end point of each VEM-Token of each segment one by one using the end point fine-tuning step, so as to align with the end point of VEM-Token at the corresponding position of the first segment, until all loop segments end.

2. The method of claim 1, wherein, The beat model further comprises a basic sub-model, specifically comprising: ST110: Using the VEM processor included in the VEM-Token vocal emotion multi-modal model, set multiple filter layers to convert the sample file and the user file into a spectral format file to separate the human singing, the accompaniment sound of the instrument, and the overtone and fluctuation of the emotion, specifically comprising: Singing filter, filter frequency range 80-1500Hz, stop band attenuation ≥48dB / oct; Accompaniment filter, filter frequency range 25-3000Hz, stop band attenuation ≥36dB / oct; Emotion overtone filter, filter frequency range 1400-12000Hz; Emotion fluctuation filter, filter frequency range 0.01-1.0Hz; oct is the rule of octave in music theory, that is, the audio frequency of each oct differs by a factor of one; ST111: The beat model formula comprises: beat 1n = [sp 1n ep 1n ] 2.3, beat 2n = [sp 2n ep 2n ] 2.4, ep 1n = sp 1(n+1) 2.5, ep 2n = sp 2(n+1) 2.6, t 1n = ep 1n - sp 1n 2.7, t 2n = ep 2n - sp 2n 2.8, BEAT 1N = [(sp 11 ep 11 )...(sp 1N ep 1N )] 2.9, BEAT 2N = [(sp 21 ep 21 )...(sp 2N ep 2N )] 2.10, Δst n = sp 1n -sp 2n 2.11, Δet n = ep 1n - ep 2n 2.12, Wherein: Equation 2.1 is the VEM-Tokem sequence of the sample file, which is equivalent to the BEAT sequence of the sample file 1N N is the total number of beats of the sample file, Equation 2.2 is the VEM-Token2 sequence of the user file, which is equivalent to the BEAT sequence of the user file 2N N is the total number of beats of the user file, Formula 2.3, 2.4 are the nth beat content in VEM-Token1 sequence, VEM-Token2 sequence respectively, n is the beat sequence number, from 1 to N, Formula 2.5, 2.6 represent the nth beat end point position, which is equal to the start point position of the nth+1 beat, sp 1n , sp 2n are the n th beat start pointer and position in the sequences VEM- Token1, VEM-Token2, respectively, ep 1n , ep 2n , respectively, the end-of-beat pointer and position in the nth beat in VEM-Token1, VEM-Token2, respectively, and also the start-of-beat pointer and position in the n-1th beat, Equations 2.7, 2.8 are the time length of beat 1n , beat 2n ​ Formula 2.9, 2.10 are the contents of VEM-Token1 sequence, VEM-Token2 sequence respectively, Formula 2.11 is the start time difference of the sample file and the user file at the n th beat; Formula 2.12 is the end time difference of the sample file and the user file at the n th beat; In formula 2.13, Δt is the minimum absolute value of all beat errors, that is, the minimum value in the absolute value of the logical or of all start time differences and all end time differences; ST112: The beat includes rhythm, and the timing unit uses the number of beats or rhythms per minute, and the time length is the number of mS; ST113: The rhythm includes uniform rhythm, which keeps the same rhythm before and after in a period of time, and mutation rhythm, which starts to mutate into another uniform rhythm after a uniform rhythm, and the error between a uniform rhythm is less than 5 mS; ST114: According to music theory, the beat includes strong, weak, and secondary strong beat types; ST115: According to music theory, the vocal file VEM-Token sequence is also divided into measures, each of which includes more than one beat, and the timing unit of the measure is determined according to the beat; ST116: According to music theory, the vocal file includes a loop section, and a loop section includes more than one measure.

3. The method of claim 2, wherein, The beat capture model includes a harmonic impact sub-model: ST120: For a vocal file including percussion, separate the spectral format file into a harmonic component spectrum graph including sustained sound and an impact component spectrum graph including transient sound; ST121: According to the impact component spectrum graph, detect the spectrum energy mutation point as a rhythm pointer candidate point; ST122: Set the energy peak threshold condition, and detect the peak point meeting the condition as the rhythm pointer candidate point; ST123: According to the rules of music theory, refer to the harmonic component spectrum graph, detect the rhythm pointer candidate point, delete the non-main rhythm candidate point, and reserve the main rhythm point; ST124: Refer to the harmonic component spectrum graph, align the main rhythm point, and generate the beat start point.

4. The method of claim 2, wherein, The beat capture model also includes a joint learning sub-model: ST130: Use a time-frequency convolution network to extract multiple spectral features, model the time sequence through a bidirectional recurrent network, and output a beat position heat map and a time signature classification probability in the output layer. According to the time signature type, weight the heat map and extract the rhythm pointer; ST131: Fuse the three channels of mel spectrum, constant Q transform spectrum, and instantaneous phase spectrum to construct a multi-channel time-frequency input, input it into a convolution recurrent neural network CRNN, and the output layer includes a first branch in parallel, which outputs the probability of the existence of a beat in each time frame, and a second branch in parallel, which outputs the time signature classification probability; ST132: According to the predicted time signature type, the weight is corrected, including: if it is 4 / 4 beat, the first and third beat probability values are multiplied by the weighting coefficient, and if it is 3 / 4 beat, the first beat probability value is multiplied by the weighting coefficient; ST133: Take the local maximum value point of the extracted corrected probability as the final rhythm pointer.

5. The method of claim 2, wherein, The beat capture model also includes a harmonic frequency hierarchical sub-model: ST140: Use a convolutional neural network to convert the sample file and the user file into multiple hierarchical spectral format files, including beat frequency, vocal fundamental frequency, harmonic fundamental frequency, vocal overtone, instrument fundamental frequency, and instrument overtone, respectively: Voice fundamental frequency is 80-1400Hz, Harmonic fundamental frequency is 100-500Hz, Voice overtone is 1.5-8kHz, Instrument fundamental frequency is 27.5-4186Hz, Instrument overtone is 4-16kHz; ST141: Using adaptive convolution network, the upper and lower limits of each sub-layer frequency are calculated to obtain the rhythm pointer.

6. The method of claim 2, wherein, The beat capture model also includes a dynamic time warping sub-model: ST150: Extracting phonemes, beats and lyric sentences in the song from sample files and user files, and positioning includes: Phoneme boundary anchor points, locating initial consonant-vowel transition points or plosive starting points by hidden Markov model as rhythm pointers; Beat strength anchor points, marking strong beats and secondary beats based on energy envelope maximum points as rhythm pointers; Sentence gas port anchor points, detecting <-40dB silent segments and spectral mutation points as rhythm pointers; The beat time error is not more than 50 milliseconds.

7. The method according to any one of claims 3 to 6, characterized in that, The beat alignment model includes a basic model: ST210: Based on the setting, for one beat 2n The alignment state of the user beat and the sample beat includes: When Δst n ≤ Δt, the start point pointer state of the n th beat of the user file and the sample file is aligned, When Δst n < - Δt, the start pointer state of the n th beat of the user file and the sample file is behind, When Δst n > Δt, the start pointer state of the n th beat of the user file and the sample file is ahead, The beats of vocal files include: concertos, ballads, strong rhythm songs, and loose rhythm songs with unclear rhythm; ST220: The value judgment range of beat error Δt includes: For songs, Δt≥10mS, For accompaniment, Δt≥5mS, For emotions, Δt≥20mS, For vocal files of strong rhythm or slow rhythm ballads, the beat error is reduced and enlarged according to the song, but the value is not greater than 20% of the corresponding beat, For loose rhythm, Δt is extended or cancelled.

8. The method of claim 7, wherein, The beat model also includes start point fine tuning and end point fine tuning: For a beat numbered n, the start point fine tuning step and the end point fine tuning step are executed in turn; ST230: For the start point pointer state of the beat, the start point fine tuning step includes: When Δt < 100mS, then translate beat left 2n with a translation time of Δt, When Δt > 100mS, then at beat 2n Add a length of Δt of decorative sound on the left, including tremolo, slide, breath, or, at beat 2n Copy the length of Δt of the content inside to extend, or, at beat 2n Clone the length of Δt of the content inside to extend; ST240: For the case where the start pointer of the beat is ahead of the state, the start fine-tuning step includes shifting right beat 2n , with a shift time of Δt, or, inside beat 2n , deleting the content of the duration of Δt to extend; ST310: in accordance with the beat after the execution of the start point fine adjustment step 2n , recalculate and calibrate the end point pointer of the beat numbered n, and calculate the end point alignment state with the beat 1n , obtain the end point alignment state, the end point lag state, and the end point lead state; ST320: For the end point pointer state of the beat, the end point fine tuning step includes: When Δt≤100mS, the end point is shifted to the left by Δt, When Δt > 100mS, then compress beat 2n The compression mode is to compress with the same frequency as the original beat 2n ​ ST330: For the end pointer of the beat state ahead, the end fine-tuning step includes inserting a trill, a slide, a breath, a rest, or, at beat 2n Internally replicate the duration of Δt to extend, or, at beat 2n Internally clone the duration of Δt to extend.

9. The method of claim 8, wherein, The beat alignment model also includes full-range alignment and verification steps: ST410: Taking each beat in the entire VEM-Token1 sequence as a working set, starting from the first beat and executing the start point fine tuning step and the end point fine tuning step in turn to complete the alignment of the VEM-Token2 sequence to the VEM-Token1 sequence; ST420: According to the start points and end points of the VEM-Token1 sequence, the alignment state of the start points and end points of the VEM-Token2 sequence is verified, and the user manually processes the unaligned beat part.

10. The method of claim 9, wherein, The beat model also includes a beat editor, which specifically includes: ST430: The beat editor provides a user interface for defining, editing, collecting sample files, recording user files, determining beats, and aligning beats; ST440: The beat editor includes functions of automatically determining beats, start point calibration, end point calibration, start point alignment, end point alignment, beat data display, and manual adjustment of beat determination, start point calibration, end point calibration, start point alignment, end point alignment, and beat data display based on the needs of vocal emotions. ST450: Change the rhythm of the VEM-Token1 sequence according to the measure or cycle segment, and make cross-measure or cross-cycle segment beat modification and beat alignment; ST460: According to the style of the vocal music file, mark the VEM-Token sequence with equal beat time intervals as uniform beat, mark the VEM-Token sequence of the vocal music file with emotional style as local uneven beat, and mark the VEM-Token sequence of the vocal music file with the style of the spread board as global uneven beat.

11. The method of claim 10, wherein, The beat capture and alignment model also includes a non-aligned personalized free beat step, which specifically includes: ST500: According to the special processing of the user for the vocal emotion, the special processing local beat in the VEM-Token2 sequence, the start point of the beat is not aligned with the start point of the corresponding beat in the corresponding VEM-Token1 sequence, and the end point of the beat is not aligned with the end point of the corresponding beat in the corresponding VEM-Token1 sequence, but for the user emotion, the start point and the end point of the special processing local beat are aligned, and the start point and the end point are aligned in the special processing local beat. In advance, lag behind; ST510: After the special processing local beat, restore the alignment of the start point and the end point of the beat; ST520: The beat editor supports editing and marking of the free beat.

12. The method of claim 11, wherein, Also includes: ST600: The beat capture and alignment model also includes a repeat alignment step; ST610: When the beat alignment model is submitted to the subsequent large model once, after the large model runs the processing result, the result is traced back to this beat alignment model again to recheck the beat alignment scheme, and if there is a deviation in the beat alignment, the operation of re-beat alignment is performed again; ST620: The step of providing a beat alignment synchronization signal to the subsequent large model application, the synchronization signal including beat number, beat type, and beat width information; ST630: According to the complete vocal music file, output the beat-aligned score including the staff, the simplified score and the beat score.

13. The method of claim 12, wherein, Also includes communication interface protocol: ST700: According to the beat alignment synchronization signal provided to the subsequent large model application, through the bidirectional communication protocol and bidirectional interface protocol for intelligent agents, together with the subsequent large model, form an AI intelligent agent Agent; ST800: According to the beat alignment synchronization signal provided to the subsequent large model application, through the bidirectional communication protocol and bidirectional interface protocol for independent AI music system, form an AI music system; ST900: The bidirectional communication interface protocol includes: the uplink data, downlink data, and handshake rules of the method for intelligent agent Agent or AI music system.

Citation Information

Patent Citations

  • Beat extraction device and beat extraction method

    CN101375327A

  • VEM-Token vocal music emotion multi-mode token song and accompaniment deep learning method

    CN120126506A