Construction method of vem-token vocal emotion multimodal magic modification model
The VEM-Token vocal emotion multimodal modification model enables automated recognition and intelligent modification of vocal emotions, overcoming the shortcomings of existing technologies in vocal emotion recognition and modification, and improving the efficiency and standardization of the modification process.
Patent Information
- Application Number
- CN202511340091.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing technologies cannot effectively identify and understand vocal emotions, lack automated and intelligent vocal modification capabilities, and rely on manual operation, which is inefficient.
The VEM-Token vocal emotion multimodal modification model is adopted. By capturing and aligning beats and combining VEM parameters, it realizes the automatic processing of singing, accompaniment and emotion. It uses supervised learning and deep learning to build a VEM library, performs vocal emotion recognition and understanding, and performs automatic modification.
It has achieved automated and intelligent recognition and modification of vocal emotions, reducing reliance on manual parameter adjustment and improving the efficiency and standardization of the modification process.
Smart Images

Figure CN120853611B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to the model construction and processing of AI intelligent agent, AI music and voice recitation, and more particularly to an innovative method of dividing tokens using music rhythm and vocal emotion multi-modal model. Specifically, when people imitate sample vocal music or sample recitation, the beat capture and beat alignment process is performed to achieve and optimize the imitation effect, and the construction method of the VEM-Token vocal emotion multi-modal modified model is constructed. BACKGROUND
[0002] Currently, in the field of artificial intelligence, the decomposition of information is still based on the natural language token division method of NLP-Token (Natural-Language-Processing Token). If it is based on the text information modality, NLP-Token has a natural advantage, and large models have learned all human books and web pages based on text and memorized them in large models. If it is for non-text information modality, such as song, emotion, music style and other forms of information modality, the current large model still seeks the text description in the book and web page learned from the past in the model memory to obtain "explanation" and "understanding" through these text descriptions, that is, it still uses NLP-Token based on text token information modality. Since we cannot guide the large model to obtain the corpus from where and cannot predict the correctness of the corpus at this time, the "illusion" of the large model cannot be avoided.
[0003] Specifically, the prior art includes:
[0004] (1) Based on manual editing and digital audio workstation (DAW) tools, which rely heavily on manual operation. Representative technologies and software include: Melodyne (Celemony): Through "note granulation" technology, it allows users to manually adjust pitch, duration, volume, vibrato and even the formant of each note. Auto-Tune (Antares Audio Technologies): Originally designed as a real-time pitch correction effect, its "graphic mode" can also be used for manual fine-tuning of pitch lines. Waves Tune / iZotope Nectar, etc.: Provide similar pitch and sound editing functions. Adobe Audition / Audacity / Logic Pro / Cubase, etc. DAW built-in tools: Provide basic functions such as compression, equalization (EQ), reverb, volume envelope editing, etc.
[0005] (2) Rule-based and Digital Signal Processing (DSP) based automation methods, which try to automate part of the correction work through pre-set algorithm rules. Representative technologies and software: One-click pitch correction plugins: such as Auto-Tune's "Auto Mode", Waves Tune Real-Time, etc. Beat alignment algorithms: such as Ableton Live's "WarpMarking" technology, which can automatically analyze and stretch audio to align with the grid. Automatic pitch correction: Most DAWs and plugins offer "one-click pitch correction" functions.
[0006] (3) Data-driven and Machine Learning (ML) based methods, which mostly focus on the local application of "timbre conversion" and "vocal synthesis".
[0007] As can be seen, based on the current traditional NLP-Token method, the multi-modal understanding of vocal music emotion from the model aspect has not been solved.
[0008] A brand new model design proposed by the inventor team for the first time, including a granted Chinese invention patent "VEM-Token vocal emotion multimodal tokenization singing and accompaniment deep learning method, CN120126506" (hereinafter referred to as "VEM-Token vocal emotion multimodal model") and an invention patent under review "VEM-Token beat capture and alignment model construction method, 202511249168.0" (hereinafter referred to as "VEM-Token beat model"), successfully divides music files with music beats as information tokens, i.e. VEM-Token (Vocal-Emotion-Multimodal Token), and explains and understands the meaning of music through vocal emotion multimodal VEM parameters. In popular terms, it is "sound into text", i.e. vocal / music generates tokens. In addition, in the first two aspects, a certain number of songs with standard style and classification are collected by human music experts, the vocal files are spectrally, the beats are detected, the spectrally vocal files are segmented into VEM-Token sequences according to the vocal beats, the lyrics, singing, accompaniment, singer emotion, accompaniment emotion, video, image, etc. Multimodal, VEM coordinate system, VEM function and VEM library are established, VEM-Token recognition is performed, and song stream and accompaniment stream are separated. According to the vocal expert, the vocal sample is scored for multimodal emotion, and the VEM parameters are obtained by supervised learning and deep learning algorithm to learn the multimodal emotion of the vocal sample. For other vocal works, it can identify vocal multimodal emotion, output lyrics, VEM-Token singing score, VEM-Token accompaniment score and VEM-Token score. The present application is a patent pool patent of CN120126506 and 202511249168.0, which can be connected to an AI system or an independently developed application system to develop a vocal intelligent agent Agent that can listen to music and recognize score. Supervised learning has obtained a VEM library.
[0009] Due to the existence of this professional and accurate VEM library learned by human experts, the possibility of "hallucination" produced by the application of the later large model almost disappears.
[0010] The VEM-Token concept mentioned in the present application is based on the basic concept and steps defined in the CN120126506 and 202511249168.0 invention patents, unless otherwise emphasized or specifically defined in the present application. Among them, the "vocal file", "music file", "song", "music" mentioned in the present application file have the same meaning unless otherwise stated.
[0011] The present application aims to be applied to the research and application of current large models, such as OpenAI, DeepSeek, Google Gemini, Kimi, a large model of a bean bag, Wenxin Yanyan, etc., which form various intelligent agents after connecting the front end, back end and some applications. The present application aims to access these large models, realize two-way communication with them, form artificial intelligence applications based on music / vocal music, and even music AI agents, to expand the application of AI and provide powerful innovation and support.
[0012] Disadvantages of prior art methods
[0013] (1) There is no modeling of vocal music emotions, and there is a lack of multi-modal quantitative representation ability of vocal music emotions.
[0014] (2) It is impossible to realize recognition and understanding based on vocal music and music emotions.
[0015] (3) It is impossible to realize automatic and intelligent vocal music modification based on emotion recognition and understanding.
[0016] (4) The source of the modification method is based on manual work, and can only rely on manual parameter tuning, which is low in efficiency and has no standardization and automation. SUMMARY
[0017] According to the deficiencies of the prior art, the present application proposes a new vocal music emotion multi-modal modification idea and method, which realizes the process of automatic modification as much as possible when the user learns to sing and imitate sample songs. The purpose and intention of the present application are achieved.
[0018] The purpose and intention of the present application are achieved by using the following technical solutions and working steps:
[0019] 1. VEM-Token modification model basic scheme implementation steps
[0020] The present application is a construction method of a VEM-Token vocal music emotion multi-modal modification model, which includes but is not limited to the following steps:
[0021] ST100: Collect sample files and user files, capture beats and VEM-Token segmentation according to the VEM-Token model, obtain VEM-Token1 sequence of the sample file and VEM-Token2 sequence of the user file respectively, align beats according to the VEM-Token1 sequence, and generate VEM-Token2 sequence for all VEM-Token2 sequences.
[0022] ST200: According to the VEM parameters included in the VEM-Token model, identify the VEM parameters of the VEM-Token1 sequence, determine the modification scheme by the user, process the VEM parameters of the VEM-Token2 sequence, and generate the VEM-Token2 sequence of the result file in the style of the VEM-Token1 sequence according to the modification scheme.
[0023] 2, Sample file and user file beat capture alignment step
[0024] On the basis of the foregoing scheme, the present application includes, but is not limited to, one or more combinations of the following beat capture and alignment steps or methods in the sample file and user file model:
[0025] ST110: For a sample file that does not include a loop segment, use beat capture to mark the start and end points of the beat, and generate all VEM-Token1 sequences.
[0026] ST120: For a sample file that includes a loop segment, from the second loop segment to the end of all loop segments, after beat capture, perform the start fine-tuning step and the end fine-tuning step to achieve beat alignment within the loop segment, and generate all VEM-Token1 sequences.
[0027] ST130: According to the VEM-Token1 sequence, perform beat capture and beat alignment steps including the start fine-tuning step and the end fine-tuning step in all VEM-Token2 sequences to process the corresponding beat, and generate VEM-Token2 sequences.
[0028] ST140: The user file specifically includes: a user file produced by the user imitating the sample file to sing, a user file produced by collecting the user's voice characteristics according to the sample file to clone, and a result file produced by mixing a locally sung user file and a locally cloned user file.
[0029] 3, VEM-Token model
[0030] On the basis of the foregoing scheme, the present application includes, but is not limited to, one or more combinations of the following vocal emotion multi-modal model combinations in the VEM-Token model:
[0031] ST210: The VEM-Token model further includes a VEM library and a VEM processor.
[0032] ST211: The VEM parameters include VEM classifications, VEM functions, and established VEM coordinate systems that record various modalities in vocal emotion.
[0033] ST212: Modalities include one or a combination of lyrics, vocals, accompaniment, narration, vocal style, music, emotional basis, accompanying instruments, video, images, and ambient sound.
[0034] ST213: The emotional basis includes one or a combination of joy, sorrow, sadness, anger, fear, disgust, surprise, calmness, longing, expectation, trust, love, hate, affection, and enmity.
[0035] ST214: Vocal style includes one or a combination of folk song singing, popular song singing, rock song singing, Western song singing, pop song singing, original song singing, and opera singing.
[0036] ST215: Voice style includes the user's inherent vocal cord and resonance cavity fundamental frequency and overtone combination when singing, speaking, or reciting, which are inherent characteristics that distinguish the user from others.
[0037] ST216: VEM classification collects various vocal files. Human vocal experts or learning algorithms evaluate the emotions of the vocal files in terms of singing and accompaniment. Supervised learning and deep learning are used to train the VEM function to obtain VEM parameters, which are then added to the VEM library.
[0038] ST220: The VEM processor provides a user interface for machine interaction, enabling the implementation of the modified solution.
[0039] 4. Modified solution model
[0040] Based on the aforementioned solution, this invention, in terms of modified solution models, includes, but is not limited to, one or more combinations of the following steps or methods:
[0041] The modification plan includes steps for modifying the vocals, accompaniment, emotional overtones, and emotional fluctuations, specifically including:
[0042] ST310: The modification plan includes vocal enhancements, specifically:
[0043] ST311: Uses a VEM processor to preprocess sample files and user files, sets a vocal filter, converts sample files and user files into spectrum format files, and separates the VEM-Token1.1 sequence of vocals in the sample files and the VEM-Token2.1 sequence of vocals in the user files.
[0044] ST312: According to the model of beat capture and beat alignment, the beat start points and beat end points of all VEM-Token1.1 sequences and all VEM-Token2.1 sequences are captured respectively, and the beat start points and beat end points of VEM-Token2.1 sequences are aligned with the beat start points and beat end points of VEM-Token1.1 sequences by using the start point fine-tuning model and the end point fine-tuning model.
[0045] ST320: The modification scheme also includes accompaniment sound modification, specifically including:
[0046] ST321: The VEM processor is used for preprocessing sample files and user files, setting accompaniment filters, converting sample files and user files into spectral format files, separating VEM-Token1.2 sequences of accompaniment sounds of sample files, and VEM-Token2.2 sequences of accompaniment sounds of user files.
[0047] ST322: According to the model of beat capture and beat alignment, the beat start points and beat end points of all VEM-Token1.2 sequences and all VEM-Token2.2 sequences are captured respectively, and the beat start points and beat end points of VEM-Token2.2 sequences are aligned with VEM-Token1.2 sequences by using the start point fine-tuning model and the end point fine-tuning model.
[0048] ST330: The modification scheme also includes emotion overtone modification, specifically including:
[0049] ST331: The VEM processor is used for preprocessing sample files and user files, setting emotion overtone filters, converting sample files and user files into spectral format files, separating VEM-Token1.3 sequences of emotion overtones of sample files, and VEM-Token2.3 sequences of emotion overtones of user files.
[0050] ST332: According to the model of beat capture and beat alignment, the beat start points and beat end points of all VEM-Token1.3 sequences and all VEM-Token2.3 sequences are captured respectively, and the beat start points and beat end points of VEM-Token2.3 sequences are aligned with VEM-Token1.3 sequences by using the start point fine-tuning model and the end point fine-tuning model.
[0051] ST340: The modification scheme also includes emotion fluctuation modification, specifically including:
[0052] ST341: Preprocessing the sample file and the user file by using the VEM processor, setting the emotional fluctuation filter, converting the sample file and the user file into a spectrum format file, and separating the VEM-Token1.4 sequence of the emotional fluctuation of the sample file and the VEM-Token2.4 sequence of the emotional fluctuation of the user file.
[0053] ST342: Capturing the beat start point and the beat end point of the VEM-Token1.4 sequence and the VEM-Token2.4 sequence respectively according to the beat capture and beat alignment model, and aligning the beat start point and the beat end point of the VEM-Token2.4 sequence with the VEM-Token1.4 sequence by using the start point fine-tuning model and the end point fine-tuning model.
[0054] 5. Learning to sing and modifying
[0055] On the basis of the foregoing scheme, the present application is a learning-to-sing-and-modifying model for a user to learn to sing a sample song file, which includes but is not limited to one or more combinations of the following steps or methods:
[0056] The modification scheme also includes a learning-to-sing-and-modifying step, which specifically includes:
[0057] ST410: The learning-to-sing-and-modifying step includes pure learning to sing, which specifically includes:
[0058] ST411: Using the VEM processor, learning to sing more than once by a user from the whole or part of the sample file in units of a measure composed of multiple beats, recording and converting the multiple groups of song VEM-Token2.1 sequences of the user file into a user file, and selecting an optimal group of song VEM-Token2.1 sequences from the multiple groups of song VEM-Token2.1 sequences as the selected song VEM-Token2.1 sequence.
[0059] ST412: Using the ST310, ST311, and ST312 steps to generate the beat-captured and aligned VEM-Token2.1 sequence from the selected song VEM-Token2.1 sequence.
[0060] ST420: The learning-to-sing-and-modifying step also includes mixed learning to sing, which specifically includes:
[0061] ST421: Setting the weights A and B for the VEM-Token1.1 sequence and the VEM-Token2.1 sequence, and calculating the mixed learning-to-sing VEM-Token2.1 sequence by using the formula A x VEM-Token1.1 sequence + B x VEM-Token2.1 sequence, where A is less than 0.3 and A + B = 1.0.
[0062] 6. Voice cloning
[0063] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following steps or methods on the basis of the model of voice cloning magic modification:
[0064] The magic modification scheme further includes the steps of voice cloning magic modification, specifically including:
[0065] ST430: The magic modification scheme further includes voice cloning magic modification, specifically including:
[0066] ST431: Query the VEM library, and if there is a VEM parameter of the customer's voice style, obtain the VEM parameter of the customer's voice style.
[0067] If there is no VEM parameter of the customer's voice style, or the existing VEM parameter of the customer's voice style does not meet the user's requirements, a segment of the user's singing recording of the practice song, a segment of the recording of speaking and reciting are collected, and the recording is converted into a frequency spectrum format file with the user's voice style by the VEM processor using the song filter, the emotion overtone filter, and the emotion fluctuation filter. According to the VEM classification, the VEM parameter of the user's voice style is obtained and stored in the VEM library.
[0068] ST432: According to the song VEM-Token1.1 sequence of the sample file, the lyrics 1 in the sample file are identified by using the speech recognition included in the VEM-Token model, and the lyrics 1 include the lyrics and the positions of the start and end points of the lyrics in the beat.
[0069] ST433: Copy the lyrics 1 of the sample file to become the lyrics 2 of the user file, and according to the VEM parameter of the user's voice style, the speech synthesizer is used to complete the voice cloning magic modification of the user according to the lyrics 2 of the user file and its start and end positions in the beat, becoming the song VEM-Token2.1 sequence of the user.
[0070] 7. Lyrics magic modification
[0071] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following steps or methods on the basis of the model of lyrics magic modification:
[0072] The magic modification scheme further includes the steps of lyrics magic modification, specifically including:
[0073] ST500: When the lyrics 2 and the lyrics 1 are inconsistent or the user needs to modify, the lyrics magic modification is performed, specifically including:
[0074] ST510: According to the semantic syntax, the lyrics 2 and the lyrics 1 are decomposed into lyrics sentences 2 and lyrics sentences 1, and the following steps are performed:
[0075] ST511: The number of words in lyrics sentence 2 is the same as that in lyrics sentence 1, lyrics score 2 is copied as lyrics score 1 according to the beat, and each word in the lyrics in lyrics sentence 2 is filled in the corresponding beat position one by one, or is manually modified by the user to become the modified lyrics score 2.
[0076] ST512: The number of words in lyrics sentence 2 is different from that in lyrics sentence 1, and the lyrics of lyrics 2 are manually modified by the user to be filled in the corresponding beat position one by one to become the modified lyrics score 2.
[0077] ST513: A voice synthesizer is used to complete the voice cloning modification of the user according to the VEM parameters of the voice style of the user, the lyrics score 2 and the start and end positions of the beat, and the voice VEM-Token2.1 sequence of the user is obtained.
[0078] 8, pitch modification, ornamentation modification, etc.
[0079] On the basis of the foregoing scheme, the present application further includes but is not limited to one or more combinations of the following steps or methods:
[0080] The modification scheme further includes the steps of pitch calibration modification, ornamentation modification, beat length modification, rhythm speed modification, beat strength modification, timbre modification, emotion modification, video modification, free modification, multiple sample modification, real-time listening of what is heard, which respectively specifically include:
[0081] ST520: Pitch calibration modification: according to the musical theory of twelve equal temperament, the fundamental frequency of the sound must be equal to the node frequency of the twelve equal temperament, and the sound frequency between the adjacent two node frequencies needs to be adjusted up or down to the node frequency:
[0082] ST530: Pre-ornamentation modification: in a beat, when the beat start point of the song VEM-Token2.1 of the user file lags behind the beat start point of the song VEM-Token1.1 of the corresponding sample file within 1 / 2 beat on the time axis, a decoration sound is used to compensate before the beat start point of the song VEM-Token2.1 of the user file to align the start point, and the decoration sound includes tremolo, glissando, extended sound, and breathing sound.
[0083] ST540: Post-ornamentation modification: in a beat, when the beat end point of the song VEM-Token2.1 of the user file leads the beat end point of the song VEM-Token1.1 of the corresponding sample file within 1 / 2 beat on the time axis, a decoration sound or a rest sound is used to compensate after the beat end point of the song VEM-Token2.1 of the user file to align the end point.
[0084] ST550: Beat length modification: When the length of the beat of the song VEM-Token2.1 of the user file does not match the length of the beat of the song VEM-Token1.1 of the corresponding sample file, a time stretching algorithm step or a grace note modification step is used to compress or extend the beat of the song VEM-Token2.1 of the user file to align with the beat of the song VEM-Token1.1 of the corresponding sample file.
[0085] ST560: Tempo modification: When the overall tempo of the user file needs to be sped up or slowed down, a time stretching algorithm step is used to compress or extend the tempo of the song and the accompaniment of the user file synchronously.
[0086] ST570: Beat strength modification: For the song VEM-Token2.1 of the user file and / or the accompaniment VEM-Token2.2 of the user file, different processing based on the base layer and the emotion layer is used according to the user's needs, wherein:
[0087] ST571: The base layer includes a step of volume dynamic processing and envelope shaping, which is adjusted above a threshold for VEM-Token2.1 and VEM-Token2.2.
[0088] ST572: The emotion layer includes intelligent dynamic control based on AI / machine learning, specifically including training a model to intelligently identify the beat, instruments in the audio, and automatically generating dynamic processing parameters according to pre-set emotion labels or target loudness curves, and the training results are stored in the VEM library.
[0089] ST580: When the length of the beat of VEM-Token2.1 changes, the length of the beat of VEM-Token2.2 needs to be verified synchronously, and when the lengths of the beats are inconsistent, time stretching is used to capture and align VEM-Token2.1 and VEM-Token2.2.
[0090] ST590: Tone modification, specifically including:
[0091] ST591: For the song VEM-Token1.1 sequence of the sample file and the song VEM-Token2.1 sequence of the user file, a filter including multiple groups of high-order harmonics of the fundamental frequency is used to decompose the frequencies of the multiple groups of high-order harmonics: F1, F2, F3, …, Fn, and the amplitudes of the high-order harmonics: A1, A2, A3, …, An, where n is the number of harmonics, and n is less than 50.
[0092] ST592: Adopting the steps of volume dynamic processing and envelope shaping, respectively amplifying or reducing the amplitude of one or more specified harmonic components in A1, A2, A3,..., An in the VEM-Token2.1 sequence to change the timbre of the user file.
[0093] ST593: Querying the AI / machine learning timbre dynamic processing parameters in the VEM library, adjusting the dynamic processing parameters to change the timbre of the user file.
[0094] ST5A0: Emotional magic modification, specifically including: according to the user's emotional magic modification requirements, querying the VEM library with the VEM-Token2.1 sequence as the independent variable, respectively adjusting the VEM parameters including emotional magic modification, vocal style, and vocal style requirements to obtain the emotional magic modification result.
[0095] ST5B0: Video magic modification, when the user file needs to be adapted to the video, adjusting the content and rhythm of the video according to the VEM parameters and rhythm to adapt to the needs of the user file.
[0096] ST5C0: Free magic modification, the user modifies the content and rhythm of VEM-Token2.1, VEM-Token2.2, and video according to the VEM parameters and one or more modalities to adapt to the needs of the user file.
[0097] ST5D0: Multi-sample magic modification, the user selects one or more sample files and respectively selects part of the parameters corresponding to part of the sample files and selects another part of the parameters corresponding to another part of the sample files to modify the content and rhythm of VEM-Token2.1, VEM-Token2.2, and video to adapt to the needs of the user file.
[0098] ST5E0: Real-time listening of what you hear is what you get, specifically including real-time listening, judging, and scoring of the modified results by the user, and submitting the modified process and modified results to the VEM-Token2 sequence to the VEM library.
[0099] 9, Member management
[0100] On the basis of the foregoing scheme, the present application further includes but is not limited to member management for users, specifically including one or more combinations of the following steps or methods:
[0101] The magic modification model also includes member management, specifically including:
[0102] ST600: According to the user's needs, apply for membership to the magic modification model, establish a membership file, and store it in the VEM library.
[0103] ST610: Member profile includes user information, sample file information, user file information, VEM parameters, voiceprint encryption, voiceprint decryption, wherein the encryption and decryption keys include member signature, member image, member video, member VEM parameters.
[0104] ST620: Member management includes forward, backward, rollback, add, delete, query, modify, store, and keep of the operation steps in the magic modification process.
[0105] ST630: Member management also includes real-time modification, real-time monitoring, real-time scoring, supervised learning, reinforcement learning, rewards, and punishments of user files, and the results are stored in the VEM library.
[0106] 10. Others
[0107] On the basis of the foregoing scheme, the magic modification model of the present application further includes but is not limited to the following steps or methods:
[0108] ST700: The magic modification model further includes application systems of mobile terminals and PC terminals, and also includes application systems of cloud mode and block chain.
[0109] ST800: The magic modification model further includes supporting hardware systems, including communication interfaces, recording modules, tuning modules, playback modules, encryption modules, decryption modules, interfaces for Douyin system, interfaces for Zhihuansheng system, and supporting AI karaoke systems.
[0110] ST900: According to the application of the subsequent large model, a synchronous signal is provided, an AI system including DeepSeek, Kimi.AI, and ChatGPT is accessed, and an AI intelligent agent is formed.
[0111] STA00: The magic modification model further includes interface protocols, provides MIDI protocols, MSC extension protocols, and OSC network protocols based on hardware and network, provides AES3 / PDIF protocols and MADI protocols based on transmission layer, and provides network audio transmission protocols such as Dante protocol, AVB / TSN protocol, and AES67 protocol.
[0112] 11. Purpose and intention of the application
[0113] The construction method of the VEM-Token vocal emotion multi-modal magic modification model has the purposes and intentions of:
[0114] "Sound generates text" is realized, and recognition and understanding based on vocal music and music emotion are realized.
[0115] An automatic and intelligent vocal AI magic modification model for creating music file emotion recognition and understanding is created, and VEM parameters are introduced.
[0116] The application discloses an automatic modification method and steps for creating a user to learn from a sample music file.
[0117] The probability of automatic parameter adjustment in the modification process is greatly improved, and standardization and automation are realized.
[0118] An innovative music / vocal music-oriented model is provided, which accesses a large model to form a music AI application agent or a dedicated application system, thereby providing strong support for expanding the application of AI.
[0119] The application realizes vocal music emotion modeling and multi-modal quantitative representation of vocal music emotion.
[0120] 12、Inventive benefits
[0121] (1) The application realizes "sound-to-text", and enables AI to recognize and understand music / sound emotion.
[0122] (2) The application solves the vocal music emotion modeling and establishes the multi-modal quantitative representation of vocal music emotion.
[0123] (3) The application realizes automatic and intelligent vocal music modification based on emotion understanding.
[0124] (4) The application realizes an efficient modification method by using automatic, semi-automatic and artificial intelligence to replace manual parameter adjustment.
[0125] (5) The application greatly reduces the "illusion" of AI operation. BRIEF DESCRIPTION OF DRAWINGS
[0126] LIST OF DRAWINGS
[0127] Figure 1 : VEM-Token modification model schematic diagram
[0128] Figure 2 : Three-dimensional space correlation emotion schematic diagram
[0129] Figure 3 : Local user operation interface schematic diagram
[0130] DETAILED DESCRIPTION OF DRAWINGS
[0131] See the specific embodiments. DETAILED DESCRIPTION
[0132] The present application is a patent pool patent of a granted Chinese invention patent "VEM-Token vocal emotion multi-modal tokenization singing and accompaniment deep learning method, CN120126506" and an invention patent "VEM-Token beat capture and alignment model construction method, 202511249168.0" which is under examination. The present application focuses on the imitation and modification of the user learning sample songs and makes further basic innovations.
[0133] The purpose and intention of the present application can be achieved by the following specific embodiments. It should be particularly noted that the specific embodiments have specific uses and industrial applicability. Therefore, the embodiments do not include all the features and steps of the present application, nor are they a limitation of the present application. The description of the claims of the present application is the summary of the invention.
[0134] This example is one of the examples of the present application.
[0135] The specific embodiments of the present application are as follows:
[0136] Innovative VEM-Token vocal emotion multi-modal modification model construction method--an AI singing learning agent and AI karaoke system
[0137] Diagram explanation
[0138] The contents of the present embodiment mainly include, but are not limited to, the following main schematic diagrams, which are: Figures 1 to 3 .
[0139] Implementation step explanation
[0140] The method steps of the present embodiment mainly include steps 1 to 10. Each of the 10 parts includes several sub-steps. Unless otherwise specified, these sub-steps are not completely required, and unless otherwise specified, the order is not essential, but is optimized and further selected by the patent implementer according to the needs of some specific tasks.
[0141] The specific work steps are as follows:
[0142] 1. VEM-Token modification model basic scheme implementation steps
[0143] The present application as a VEM-Token vocal emotion multi-modal modification model construction method includes, but is not limited to, the following steps:
[0144] ST100: Collect sample files and user files, capture and segment VEM-Tokens according to the VEM-Token model, obtain VEM-Token1 sequence of sample files and VEM-Token2 sequence of user files, respectively, and generate VEM-Token2 sequence by beat alignment according to VEM-Token1 sequence.
[0145] ST200: Identify VEM parameters of VEM-Token1 sequence according to VEM parameters included in VEM-Token model, determine modification scheme by user, process VEM parameters of VEM-Token2 sequence, and generate VEM-Token2 sequence of result files that conforms to modification scheme and style of VEM-Token1 sequence.
[0146] Among them, the model of VEM-Token vocal emotion multi-modal refers to the model in "VEM-Token vocal emotion multi-modal tokenization singing and accompaniment deep learning method, CN120126506", which specifically includes steps 1-4:
[0147] (1) Record emotions by using one or more modalities, mark vocal emotion multi-modal as VEM, and construct VEM classification, VEM coordinate system, VEM function and VEM library. Vocal emotion includes one or a combination of joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, emotion, and enmity. Multi-modal includes one or a combination of lyrics, singing, accompaniment, vocal style, music, emotion basis, accompaniment instrument, video, and image. VEM coordinate system includes a coordinate axis system established according to independent emotions, opposite emotion pairs, and related opposite emotion groups.
[0148] (2) Collect vocal samples according to VEM classification, and perform emotion evaluation on vocal samples in singing and accompaniment by human vocal experts. Use supervised learning and deep learning to train VEM function to obtain VEM parameters and add them to VEM library.
[0149] (3) Use VEM processor to beat mark vocal files and separate singing stream and accompaniment stream. According to the beat, VEM-Token is segmented to convert singing stream into VEM-Token1 sequence and accompaniment stream into VEM-Token2 sequence, and add them to the preprocessing library.
[0150] (4) Use deep learning to generate lyrics score, VEM-Token singing score, VEM-Token accompaniment score, and VEM-Token score.
[0151] The model of VEM-Token beat capture and beat alignment refers to the model in the VEM-Token beat capture and alignment model construction method, 202511249168.0, and specifically includes steps 5-7.
[0152] (5) For a vocal file, a beat model including beat capture and beat alignment is set according to a VEM-Token vocal emotion multi-modal model, the beats of the vocal file are captured, and the vocal file is divided into VEM-Token sequences according to the beats, and the positions of the start of the beat and the end of the beat in each VEM-Token are marked;
[0153] (6) A start point alignment model is set, including:
[0154] The sample file included in the vocal file and the user file sung by the user by imitating the sample file are divided into VEM-Token1 sequences and VEM-Token2 sequences, respectively, and the start points of each VEM-Token1 are used to adjust the start points of the corresponding VEM-Token2 one by one using the start point fine-tuning step, so that the start points of the corresponding VEM-Token1 are aligned.
[0155] For each segment of the loop segment included in the vocal file, starting from the second segment, the start points of each VEM-Token of each segment are adjusted one by one using the start point fine-tuning step with the first segment as a reference, so that the start points of the VEM-Tokens corresponding to the positions of the first segment are aligned, until all the loop segments end.
[0156] (7) A terminal point alignment model is set, including:
[0157] According to the end points of each VEM-Token1, the end points of the corresponding VEM-Token2 are adjusted one by one using the end point fine-tuning step, so that the end points of the corresponding VEM-Token1 are aligned.
[0158] For each segment of the loop segment, starting from the second segment, the end points of each VEM-Token of each segment are adjusted one by one using the end point fine-tuning step with the first segment as a reference, so that the end points of the VEM-Tokens corresponding to the positions of the first segment are aligned, until all the loop segments end.
[0159] In the present application, the modified model is an algorithm for the user file to imitate the sample file.
[0160] It should be noted here that the sample file includes more than one.
[0161] If the user only needs to imitate a single sample file, which is only one song, only the VEM parameters of this song are needed. For example, the user only wants to imitate the song "West Sea Love Song" by Lao Lang, so the sample file is the version of this song by Lao Lang, and all the VEM parameters of this song are selected.
[0162] If the user needs to select different VEM parameters in two sample files, then there are two sample files. For example, the user likes to imitate part of "West Sea Love Song" by Lao Lang and part of "West Sea Love Song" by Ya Nan, so the VEM parameters of the two singers' performances of "West Sea Love Song" need to be selected.
[0163] Figure 1 : VEM-Token magic model diagram
[0164] In Figure 1 , one or more sample files, user files are input to the VEM processor, and are sent to the VEM library query by path 101. If the VEM library is found to store the sample file or even the user file, the VEM parameters of the file, and / or VEM-Token1 sequence / VEM-Token2 sequence, are sent to the beat capture and alignment module by path 103. According to the user's needs, the VEM parameters and VEM-Token1 sequence of the sample file can also be sent to the VEM-Token magic matrix by path 102. If the sample file does not exist in the VEM parameter library, the VEM processor is used to analyze the VEM parameters of the sample file and / or user file.
[0165] In the beat capture and alignment module, data from the VEM processor includes at least the VEM-Token2 sequence of the user file, and also includes the VEM-Token1 sequence of the sample file by path 103, or from the VEM processor including the VEM-Token1 sequence of the sample file and the VEM-Token2 sequence of the user file, performs the beat capture and beat alignment steps. Then sent to the multi-modal learning analysis module, further decomposed into VEM-Token1.1, VEM-Token1.2 sequence, VEM-Token2.1 sequence, VEM-Token2.2 sequence of the sample file and user file song, accompaniment sound. In addition, according to the user's needs, in terms of emotional overtones and emotional fluctuations, continue to decompose into sample emotional overtones VEM-Token1.3, sample emotional fluctuations VEM-Token1.4, and user file emotional overtones VEM-Token2.3, sample emotional fluctuations VEM-Token2.4.
[0166] In Figure 1In the present application, in order to facilitate the graphical expression, the VEM-Token modification matrix is set up, and all the modification items in the present application are included in the VEM-Token modification matrix. The user of the present patent should understand that this is only a graphical description, and is not necessarily a module or hardware, and is not a limitation of the present application.
[0167] The information source of the VEM-Token modification matrix includes two paths, which are:
[0168] All the information is from the multi-module learning analysis module, including the VEM-Token1.1, VEM-Token1.2, VEM-Token1.3, VEM-Token1.4 of the sample file, and the VEM-Token2.1, VEM-Token2.2, VEM-Token2.3, VEM-Token2.4 of the user file. In this case, the content of the sample file is not stored in the VEM library.
[0169] All the information is from the multi-module learning analysis module and the VEM library through the 104 path. The former is the content of the user file, and the latter is the content of the sample file. Here, the content of the sample file has been stored in the VEM library.
[0170] The information output of the VEM-Token modification matrix includes the modification result of the user file. According to the user's will, the modification result of the user file is stored in the VEM library.
[0171] 2, sample file and user file beat capture alignment step
[0172] On the basis of the foregoing basic scheme, the present application includes but is not limited to one or more combinations of the following beat capture and alignment steps or methods in the sample file and user file model:
[0173] ST110: For the sample file not including the loop segment, the beat capture is used to mark the start and end points of the beat, and all VEM-Token1 sequences are generated.
[0174] ST120: For the sample file including the loop segment, from the second loop segment to the end of all loop segments, after the beat capture, the start fine-tuning step and the end fine-tuning step are performed to realize the beat alignment inside the loop segment, and all VEM-Token1 sequences are generated.
[0175] ST130: According to the VEM-Token1 sequence, in all VEM-Token2 sequences, the beat capture and beat alignment steps including the start fine-tuning step and the end fine-tuning step are performed to process the corresponding beat, and the VEM-Token2 sequence is generated.
[0176] ST140: The user file specifically includes: a user file produced by the user imitating the sample file, a user file produced by collecting the user's voice characteristics according to the sample file, a result file produced by mixing a user file produced by local singing and a user file produced by local cloning.
[0177] It should be emphasized here that the loop section often appears in a song, such as a three-section loop. In general, the beats in the loop section are aligned, so the start point fine-tuning and the end point fine-tuning need to be performed. However, with the mood processing of the singer or the style of the song, the beats of individual measures in the loop section are not absolutely aligned. In this case, the user of the present patent needs to process the song differently.
[0178] 3、VEM-Token model
[0179] On the basis of the foregoing scheme, the present application includes but is not limited to one or more of the following steps or methods of a vocal music emotion multi-modal model combination in the VEM-Token model:
[0180] ST210: The VEM-Token model further includes a VEM library and a VEM processor.
[0181] ST211: The VEM parameters include VEM classification, VEM function, and established VEM coordinate system recording various modes of vocal music emotion.
[0182] ST212: The mode includes one or a combination of lyrics, singing, accompaniment, commentary, vocal music style, music, emotion basis, accompaniment instrument, video, image, and environmental sound.
[0183] ST213: The emotion basis includes one or a combination of joy, sadness, sorrow, anger, fear, disgust, surprise, calmness, longing, expectation, trust, love, hate, emotion, and hatred.
[0184] ST214: The vocal music style includes one or a combination of national song singing method, popular song singing method, rock song singing method, western song singing method, pop song singing method, original song singing method, and opera singing method.
[0185] ST215: The vocal style includes the combination of fundamental frequency and overtones emitted by the vocal cords and resonance cavity when the user sings, speaks, and recites, which is the inherent characteristic of the user that distinguishes him from others.
[0186] ST216: The VEM classification collects a variety of vocal files, and the vocal experts or learning algorithms of human beings judge the emotions of the vocal files in singing and accompaniment, adopt supervised learning and deep learning, train the VEM function to obtain the VEM parameters, and add them to the VEM library.
[0187] ST220: The VEM processor provides an operation interface for user and machine interaction, and completes the specific implementation of the magic modification scheme.
[0188] It should be pointed out here that the VEM-Token model is a model of vocal emotion multi-modal tokens, which is different from the existing NLP-Token natural language processing tokens. The NLP-Token is based on the direct interpretation of words and Chinese characters. The VEM-Token is a new interpretation of music beats as tokens. At the same time, the VEM-Token is also associated with the VEM parameters of vocal emotion multi-modal, which contains vector data of various expression modes, such as love, hatred, emotion, and enmity, as well as their direction and scale. In addition, the model also includes the VEM library of supervised learning of typical songs by human vocal experts. Therefore, in this case, the VEM parameters of the songs in the VEM library have high accuracy, and the results of the operation also have high reliability, and the probability of "hallucination" like the NLP-Token model is also very low.
[0189] 4. Magic modification scheme model
[0190] On the basis of the foregoing scheme, the present application in the aspect of magic modification scheme model includes but is not limited to one or more combinations of the following steps or methods:
[0191] The magic modification scheme includes the steps of singing modification, accompaniment sound modification, emotion overtone modification, and emotion fluctuation modification, specifically including:
[0192] ST310: The magic modification scheme includes singing modification, specifically including:
[0193] ST311: The VEM processor is used for pre-processing sample files and user files, setting a singing filter, converting the sample files and user files into frequency spectrum format files, and separating the VEM-Token1.1 sequence of the singing of the sample files and the VEM-Token2.1 sequence of the singing of the user files.
[0194] ST312: According to the beat capture and beat alignment model, the beat starting points and beat ending points of all VEM-Token1.1 sequences and all VEM-Token2.1 sequences are captured respectively, and the starting point fine-tuning model and the ending point fine-tuning model are adopted to align the beat starting points and beat ending points of the VEM-Token2.1 sequence with the beat starting points and beat ending points of the VEM-Token1.1 sequence.
[0195] ST320: The modification scheme also includes accompaniment sound modification, specifically including:
[0196] ST321: Using the VEM processor to preprocess the sample file and the user file, setting the accompaniment filter, converting the sample file and the user file into a spectrum format file, separating the VEM-Token1.2 sequence of the accompaniment sound of the sample file and the VEM-Token2.2 sequence of the accompaniment sound of the user file.
[0197] ST322: According to the beat capture and beat alignment model, respectively capture the beat start point and the beat end point of all VEM-Token1.2 sequences and all VEM-Token2.2 sequences, and use the start point fine-tuning model and the end point fine-tuning model to align the beat start point and the beat end point of the VEM-Token2.2 sequence with the VEM-Token1.2 sequence.
[0198] ST330: The modification scheme also includes emotion overtone modification, specifically including:
[0199] ST331: Using the VEM processor to preprocess the sample file and the user file, setting the emotion overtone filter, converting the sample file and the user file into a spectrum format file, separating the VEM-Token1.3 sequence of the emotion overtone of the sample file and the VEM-Token2.3 sequence of the emotion overtone of the user file.
[0200] ST332: According to the beat capture and beat alignment model, respectively capture the beat start point and the beat end point of all VEM-Token1.3 sequences and all VEM-Token2.3 sequences, and use the start point fine-tuning model and the end point fine-tuning model to align the beat start point and the beat end point of the VEM-Token2.3 sequence with the VEM-Token1.3 sequence.
[0201] ST340: The modification scheme also includes emotion fluctuation modification, specifically including:
[0202] ST341: Using the VEM processor to preprocess the sample file and the user file, setting the emotion fluctuation filter, converting the sample file and the user file into a spectrum format file, separating the VEM-Token1.4 sequence of the emotion fluctuation of the sample file and the VEM-Token2.4 sequence of the emotion fluctuation of the user file.
[0203] ST342: According to the model of beat capture and beat alignment, the beat start point and the beat end point of the whole VEM-Token1.4 sequence and the whole VEM-Token2.4 sequence are captured respectively, and the beat start point and the beat end point of the VEM-Token2.4 sequence are aligned with the VEM-Token1.4 sequence by using the start point fine-tuning model and the end point fine-tuning model.
[0204] It should be pointed out here that the four modification schemes of singing voice modification, accompaniment voice modification, emotion overtone modification, and emotion fluctuation modification are the basic modification schemes of the sample song file facing the user's learning file, one of the focuses is the capture and alignment of the beats of the two. Secondly, the four modification schemes can be combined, for example, in many cases, the accompaniment voice can not need to be modified, and the accompaniment track of the sample file can be directly selected. It is particularly emphasized that in the beat fine-tuning, after each start point fine-tuning, it is recommended to do an end point fine-tuning to ensure that the length of the beat is fixed.
[0205] 5, Learning modification
[0206] On the basis of the foregoing scheme, the present application on the user's learning modification model of the sample song file includes but is not limited to one or more combinations of the following steps or methods:
[0207] The modification scheme also includes a learning modification step, specifically including:
[0208] ST410: The learning modification includes pure learning, specifically including:
[0209] ST411: Using a VEM processor, according to all or part of the sample file, a measure composed of multiple beats is taken as a unit, the user learns more than once and records and converts into multiple groups of singing VEM-Token2.1 sequences of the user's file, and then the user selects the most optimal group of singing VEM-Token2.1 sequences from the corresponding singing VEM-Token1.1 sequences of the sample file, as the selected singing VEM-Token2.1 sequences.
[0210] ST412: Using the steps of ST310, ST311, and ST312, the selected singing VEM-Token2.1 sequences are generated into beat-captured and beat-aligned VEM-Token2.1 sequences.
[0211] ST420: The learning modification also includes mixed learning, specifically including:
[0212] ST421: Set the weights A and B for the VEM-Token1.1 sequence and the VEM-Token2.1 sequence, and calculate the mixed VEM-Token2.1 sequence by using the formula A x VEM-Token1.1 sequence + B x VEM-Token2.1 sequence, wherein A is less than 0.3, and A + B = 1.0.
[0213] The singing magic modification is the first step for a user to imitate a sample file. In general, the singing of each sentence and each beat needs to be repeatedly learned and recorded many times, and then the most satisfactory one is selected. Next, the beat start point and the beat end point of the VEM-Token2.1 sequence are aligned with the beat start point and the beat end point of the VEM-Token1.1 sequence for the selected recording.
[0214] Regarding the mixed singing, after the beat capture and alignment, the data of the VEM-Token1.1 sequence of the corresponding beat of the sample file is discounted by, for example, 5%, and the data of the VEM-Token2.1 sequence of the user's singing result is discounted by 95%, and after synthesis, it is used as the final user singing file. The purpose of this step is to include a small part of the sample song file in the user's singing result, so as to make the magic modification closer to the sample song. It should be noted here that whether 5% or more is selected needs to be determined in the listening, and the proportion cannot be too large, otherwise the effect will be poor due to the large difference between the sample and the user.
[0215] 6, Voice cloning
[0216] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following steps or methods on the voice cloning magic modification model:
[0217] The magic modification scheme also includes a voice cloning magic modification step, which specifically includes:
[0218] ST430: The magic modification scheme also includes voice cloning magic modification, which specifically includes:
[0219] ST431: Query the VEM library, and if the VEM parameters of the customer's voice style exist, obtain the VEM parameters of the customer's voice style.
[0220] If the VEM parameters of the customer's voice style do not exist, or the existing VEM parameters of the customer's voice style do not meet the user's requirements, a singing recording of a user's practice song, a speaking and reciting recording are collected, the recording is converted into a frequency spectrum format file with the user's voice style by using a singing filter, an emotion overtone filter, and an emotion fluctuation filter by a VEM processor, the VEM parameters of the user's voice style are obtained according to VEM classification, and the VEM parameters are stored in the VEM library.
[0221] ST432: According to the song VEM-Token1.1 sequence of the sample file, using the speech recognition of the VEM-Token model, the lyrics spectrum 1 in the sample file is recognized, and the lyrics spectrum 1 includes the lyrics and the start and end positions of the lyrics in the beat.
[0222] ST433: Copy the lyrics spectrum 1 of the sample file into the lyrics spectrum 2 of the user file, according to the VEM parameters of the user's voice style, using the speech synthesizer, and according to the lyrics spectrum 2 of the user file and its start and end positions in the beat, the user's voice cloning magic modification is completed word by word, becoming the user's song VEM-Token2.1 sequence.
[0223] The voice of a person is like a person's fingerprint, which has its own personalized characteristics, which mainly reflects the structural characteristics of the vocal tract of the person's voice, which is reflected in the acoustic spectrum format file, that is, the sound fundamental frequency, the type of overtones and their amplitudes, which are included in the VEM parameters, and analyzing these can find the characteristics of the voice.
[0224] From the principle, as long as the VEM library has the user's voice VEM parameters, according to the song VEM-Token1.1 sequence of the sample file, plus the user's lyrics spectrum 2, using the speech synthesizer, the user's song VEM-Token2.1 sequence can be constructed, and the user's voice cloning magic modification can be realized. Even if the VEM library does not have the VEM parameters of the user, the VEM processor can collect the voice characteristics of the user and convert them into VEM parameters, so that the voice cloning magic modification scheme can be realized.
[0225] 7. Lyrics magic modification
[0226] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following model steps or methods on the model of lyrics magic modification:
[0227] The magic modification scheme also includes lyrics magic modification steps, specifically including:
[0228] ST500: When the lyrics spectrum 2 is inconsistent with the lyrics spectrum 1 or the user needs to modify, the lyrics magic modification is performed, specifically including:
[0229] ST510: According to the semantic syntax, the lyrics spectrum 2 and the lyrics spectrum 1 are decomposed into lyrics sentences 2 and lyrics sentences 1, and the following steps are performed:
[0230] ST511: The number of words in the lyrics sentence 2 and the lyrics sentence 1 is the same, the lyrics spectrum 2 is copied to the lyrics spectrum 1 according to the beat, and each word in the lyrics in the lyrics sentence 2 is filled in the corresponding beat position, or manually modified by the user, becoming the magic modified lyrics spectrum 2.
[0231] ST512: The number of words of lyric sentence 2 and lyric sentence 1 is different, and is manually modified by a user, the lyrics of lyric 2 are filled in blanks one by one to corresponding beat positions to become a magic modified lyrics score 2.
[0232] ST513: A voice synthesizer is adopted, a user's voice cloning magic modification is completed according to the VEM parameters of the user's voice style, the lyrics score 2 and the start and end positions of the beat, and the user's voice VEM-Token2.1 sequence is obtained.
[0233] Lyric modification is a common magic modification item in the magic modification model, and is also a common magic modification item of a user for a sample song, usually, the modification of the lyrics score is mostly a partial modification of the original sample file, and the meaning of the lyrics, the word filling after the tone and unaccented are comprehensively considered, instead of improvisational modification.
[0234] 8, pitch modification, ornamentation modification, etc.
[0235] On the basis of the foregoing scheme, the present application further includes but is not limited to the following one or more combined steps or methods:
[0236] The magic modification scheme further includes pitch calibration modification, ornamentation modification, beat length modification, rhythm speed modification, beat strength modification, timbre modification, emotion modification, video modification, free modification, multiple sample modification, real-time listening steps of what is heard is what is obtained, and respectively specifically includes:
[0237] ST520: Pitch calibration modification: according to the musical theory of twelve equal temperament, the fundamental frequency of the sound must be equal to the node frequency of the twelve equal temperament, and the sound frequency between the adjacent two node frequencies needs to be adjusted up or down to the node frequency.
[0238] Pitch calibration is very important for a non-professional singer. In music theory, the pitch is the frequency of the fundamental tone, and the pitch is also called the pitch. Professional singers basically pass the pitch due to long-term training. The pitch calibration modification here can include two aspects, one is automatic calibration of all pitch, and the other is beat calibration selected by the user.
[0239] ST530: Front ornamentation modification: when the beat start point of the song VEM-Token2.1 of the user file and the beat start point of the song VEM-Token1.1 of the corresponding sample file fall within 1 / 2 beat on the time axis, the ornamentation is used to compensate before the beat start point of the song VEM-Token2.1 of the user file to align the start point, and the ornamentation includes tremolo, glissando, extended sound, and breathing sound.
[0240] ST540: Post-decorative note modification: in one beat, when the end of the beat of the song VEM-Token2.1 of the user file and the end of the beat of the song VEM-Token1.1 of the corresponding sample file are within 1 / 2 beat in time axis, a decorative note or rest note is used to compensate after the end of the beat of the song VEM-Token2.1 of the user file, so that the end is aligned.
[0241] Regarding the selection of the type of pre / post-decorative note, the length of the compensation time difference, the mood VEM parameter of the beat, and other specific circumstances are generally considered, and the user selects according to his own understanding of the song.
[0242] ST550: Beat length modification: when the beat of the song VEM-Token2.1 of the user file and the beat of the song VEM-Token1.1 of the corresponding sample file do not match in length, a time stretching algorithm step or a decorative note modification step is used to compress or extend the beat of the song VEM-Token2.1 of the user file, so as to align with the beat of the song VEM-Token1.1 of the corresponding sample file.
[0243] ST560: Rhythm modification: when the overall rhythm of the user file needs to be accelerated or slowed down, a time stretching algorithm step is used to synchronously compress or extend the rhythm of the song and accompaniment of the user file.
[0244] ST570: Beat strength modification: for the song VEM-Token2.1 of the user file and / or the accompaniment VEM-Token2.2 of the user file, different processing based on the base layer and the mood layer is used according to the user's needs, wherein:
[0245] ST571: The base layer includes steps of volume dynamic processing and envelope shaping, which are adjusted above the threshold for VEM-Token2.1 and VEM-Token2.2.
[0246] ST572: The mood layer includes intelligent dynamic control based on AI / machine learning, which specifically includes training a model to intelligently identify the beat, instruments in the audio, and automatically generating dynamic processing parameters according to the preset mood label or target loudness curve, and the training result is stored in the VEM library.
[0247] ST580: When the length of the beat of VEM-Token2.1 changes, the length of the beat of VEM-Token2.2 needs to be synchronously verified, and when the lengths of the beats are inconsistent, time stretching is needed to capture and align VEM-Token2.1 and VEM-Token2.2.
[0248] Tempo and beat are the basic physical quantities in music, and the most basic elements of human invention of music. Because, music is originated from the rhythm of the number of collective labor, and developed from the synchronized dance steps to express happiness. Therefore, the modification of tempo and beat is very important.
[0249] It should be noted that the time stretching algorithm is to change the length of the audio in time, but not to change the pitch, not to change the harmonic structure of the sound. The time stretching algorithm specifically includes but is not limited to the following three:
[0250] (1) Time domain algorithm, SOLA (Synchronous Overlap-Add) and its variants (e.g., WSOLA, SOLA-FS),
[0251] The basic principle is: divide the input signal into short segments with overlap, time scaling (by compressing or expanding the overlap area between segments) for each segment, find the best cross point (based on waveform similarity) to minimize the distortion when splicing, and recombine the processed segments into the output signal.
[0252] The basic characteristics are: advantages: fast calculation speed, suitable for real-time, low resource applications (such as old-fashioned tape simulators, simple pitch shifter), disadvantages: poor processing ability for transient (such as drum points, sound heads), easy to produce "click" sound and reverberation, music processing sound quality is not ideal, representative application: early telephone voice message speed control, some simple audio plug-ins.
[0253] (2) Frequency domain algorithm (parametric / phase vocoder), which is the most mainstream and widely used method at present, based on Short-Time Fourier Transform (STFT). Mainly includes but not limited to:
[0254] Phase Vocoder, advantages, much better sound quality than time domain methods, especially suitable for harmonic-rich audio (such as vocals, piano). Disadvantages, will produce a unique "robot voice" or "reverberation" artifacts, the processing of impact audio (such as military drums) is still not perfect.
[0255] Enhanced phase vocoder based on transient processing. Characteristics, greatly improve the processing quality of drums, bass and other music. Representative, iZotope Radius, Serato Pitch 'n Time, Ableton Live "Complex" and "ComplexPro" mode.
[0256] (3) Data-driven algorithm based on machine learning / deep learning, which is the current frontier research direction, aims to generate more natural time stretching results by learning a large amount of data. Mainly includes but not limited to:
[0257] The method based on the generation model, the principle is that the model learns the latent distribution of the audio signal. Given an input audio, the model generates the most natural and most matched output audio according to the target time length. The advantage is great potential, and theoretically it can produce the most natural and least artifact results, and even "imagine" reasonable details to fill in the stretched time. The disadvantage is that it needs a huge amount of data for training; the consumption of computing resources is huge, and it is difficult to realize real-time; the model may produce uncontrollable hallucinations (artifacts).
[0258] The method based on neural vocoder, the principle is not to directly process the waveform, but to extract the intermediate representation of the audio (such as mel spectrogram, F0 pitch, harmonic information). The advantage is that the sound quality is usually much better than traditional phase vocoder, because neural vocoder has been trained on a large amount of high-quality audio, and can reconstruct more natural sound texture. The disadvantage is that it depends on the quality of neural vocoder, and is usually computationally intensive.
[0259] ST590: Timbre modification, specifically including:
[0260] ST591: For the song VEM-Token1.1 sequence of the sample file and the song VEM-Token2.1 sequence of the user file, a filter including multiple groups of high-order harmonics of the fundamental frequency is used to decompose the frequencies of the multiple groups of high-order harmonics: F1, F2, F3, …, Fn, and the amplitudes of the overtones of the high-order harmonics: A1, A2, A3, …, An, where n is the number of harmonics, and n is less than 50.
[0261] ST592: Use the steps of volume dynamic processing and envelope shaping to amplify or reduce the amplitude of one or more specified harmonic components in A1, A2, A3, …, An in the VEM-Token2.1 sequence to change the timbre of the user file.
[0262] ST593: Query the AI / machine learning timbre dynamic processing parameters in the VEM library and adjust the dynamic processing parameters to change the timbre of the user file.
[0263] The timbre here is mainly the individualized characteristics of the user's own pronunciation, similar to the voiceprint mentioned above. It has its individualized characteristics, which mainly reflect the structural characteristics of the person's vocal resonance cavity, which is reflected in the acoustic spectrum format file as the sound fundamental frequency, the type of overtones, and their amplitudes. These are included in the VEM parameters, and analyzing these can find the characteristics of the voice.
[0264] In principle, as long as the VEM library has the user's VEM parameters, according to the sample file song VEM-Token1.1 sequence, the user's song VEM-Token2.1 sequence can be constructed, and the user's timbre magic modification can be realized. Even if the VEM library does not have the user's VEM parameters, the user's voice characteristics can be collected by the VEM processor and converted into VEM parameters to realize the timbre magic modification scheme.
[0265] ST5A0: emotional modification, specifically including: according to the user's emotional modification requirements, querying the VEM library, taking the VEM-Token2.1 sequence as the independent variable, adjusting the VEM parameters including emotional modification, vocal style, and vocal style requirements, respectively, to obtain the emotional modification result.
[0266] Emotional modification is the focus of the present application. Based on the VEM-Token model, emotions have various classifications-VEM classifications. Emotions are recorded as VEM parameters according to the degree, and emotional points are included in the corresponding coordinates.
[0267] For example, the opposite and related emotional elements "love-hate, love-hate, love-hate" are represented by X, Y, and Z orthogonal three-dimensional spherical coordinate systems, respectively. The coordinate intervals are X axis "love-hate" -100% to 100%, Y axis "joy-sorrow" -100% to 100%, and Z axis "love-hate" -100% to 100%. The three coordinate axes are related to each other in an orthogonal manner, as shown in Figure 2 For example, the emotional parameters recorded by the emotional point E (X, Y, Z) of a certain beat are (80%, -20%, 50%). This method accurately records emotions.
[0268] It is particularly reminded that the user of the present application needs to pay attention to: the VEM coordinate system includes but is not limited to:
[0269] Independent emotions include emotions that are independent of each other and have no association. A single one-dimensional coordinate axis is constructed, with the lowest point of the emotion as the coordinate 0 point and the highest point of the emotion as the maximum coordinate point.
[0270] Conversely, the emotional pair includes two emotions that are opposite to each other. A two-way one-dimensional coordinate axis is constructed, with the midpoint of the opposite emotional pair as the coordinate 0 point, the highest point of the positive emotion as the positive maximum coordinate point, and the highest point of the negative emotion as the negative maximum coordinate point.
[0271] The opposite emotion groups are associated, including a group of more than one pair of opposite emotions associated with each other, and the coordinate 0 point is aligned, and the two-way one-dimensional coordinate axis is super orthogonal, and the super plane is divided, and the positive emotions are placed on the same side of the super plane, and the negative emotions are placed on the opposite side of the super plane, and the super orthogonal coordinate system is constructed, and the highest point of the respective positive emotion is the super orthogonal coordinate positive maximum point, and the highest point of the respective negative emotion is the super orthogonal coordinate negative maximum point.
[0272] It should be noted that the independent emotion, the opposite emotion pair and the associated opposite emotion group are sometimes relative, and the modal classification is not always the same for different styles of vocal music works and different cultural backgrounds of vocal music works. Therefore, the independent emotion, the opposite emotion pair and the associated opposite emotion group are all referred to a vocal music work in extreme cases.
[0273] It should be emphasized that here "super orthogonal coordinate system" refers to the arrangement of more than 3-dimensional right-angle coordinate axes in the same system, such as 4-dimensional, 5-dimensional or more, which cannot be directly drawn on paper, so it is described and recorded by means of mathematical hyperspace. In addition, the mutual association of these super dimensions is not only orthogonal, but also may only intersect in mathematics, rather than being orthogonal, and may even be a curved coordinate based on Riemann geometry.
[0274] Figure 3 It is a local user operation interface schematic diagram. In the figure, the three ribbons respectively represent the adjustment interval interfaces of the three groups of emotion pairs, and the green circles on the ribbons are the adjustment sliding cursors of the emotion VEM parameters. Pulling the green circle cursors on the ribbons left and right can adjust the emotion values of the three groups of emotion pairs "love-hate", "joy-sorrow" and "emotion-enmity". The design of the human-computer interaction interface includes but is not limited to one-dimensional, two-dimensional screen, three-dimensional space and even multi-dimensional space interaction interfaces, and the VEM parameters include but are not limited to dragging cursor type, numerical input type, color editing type, and also include the real-time monitoring mode of "what you see is what you hear". The editing results and editing methods are stored in the VEM library.
[0275] ST5B0: video modification, when the user file needs to adapt to the video, the content and rhythm of the video picture are adjusted according to the VEM parameters and the rhythm to adapt to the needs of the user file.
[0276] Since the present application is based on the VEM-Token model, which can be understood as the "sound-to-text" mode, here the subsequent "text-to-image" model can be connected, the text information token generated by "sound-to-text" drives "text-to-image" to generate static images and dynamic videos, and the user can also manually connect videos to generate video modification.
[0277] ST5C0: Free magic, users according to VEM parameters and more than one mode, magic VEM-Token2.1, VEM-Token2.2 and video screen content and rhythm, to adapt to the needs of user files.
[0278] Here, the free magic, including by users according to their own wishes, ideas to realize for user song, user accompaniment and video screen magic, magic way includes manual and semi-automatic.
[0279] ST5D0: Multiple sample magic, users select more than one sample file, and respectively in VEM parameters to select part of the parameters corresponding to part of the sample file, select another part of the parameters corresponding to another part of the sample file, magic VEM-Token2.1, VEM-Token2.2 and video screen content and rhythm, to adapt to the needs of user files.
[0280] Here, the multi-sample magic, is based on more than one sample file to realize the magic of user files, for example, the song of the user song refers to a folk song style sample file and a more elegant tenor style sample file, in addition, the accompaniment refers to the sample file of the national musical instrument accompaniment, that is, referring to 3 sample files, to realize multi-sample magic.
[0281] ST5E0: Real-time listening, including real-time listening, judging, scoring by users on the magic results, and generating VEM-Token2 sequence of the magic process and the magic results file, and submitting to the VEM library.
[0282] Real-time listening is an important function, which is beneficial to improve the efficiency of magic, it is in the user to edit the magic, real-time listening and watching the magic results, achieve "what you hear is what you get" effect.
[0283] Here, the user of the present application needs to be reminded: the design of man-machine interface includes but is not limited to one-dimensional, two-dimensional screen, three-dimensional space and even multi-dimensional space interaction interface, VEM parameters include but are not limited to drag cursor type, numerical input type, color editing type, and also include "what you see is what you hear" real-time listening mode. The editing results and editing methods are stored in the VEM library.
[0284] 9, Member management
[0285] On the basis of the foregoing scheme, the present application also includes but is not limited to the member management of users, specifically including the following one or more combinations of steps or methods:
[0286] The magic model also includes member management, specifically including:
[0287] ST600: According to the user's needs, apply for membership to the modified model, establish a membership file, and store it in the VEM library.
[0288] ST610: The membership file includes user information, sample file information, user file information, VEM parameters, voiceprint encryption, and voiceprint decryption, wherein the encryption and decryption keys include member signature, member image, member video, and member VEM parameters.
[0289] ST620: Member management includes forward, backward, rollback, add, delete, query, modify, store, and keep in the operation steps of the modification process.
[0290] ST630: Member management also includes real-time modification, real-time monitoring, real-time scoring, supervised learning, reinforcement learning, rewards, and punishments of user files, and the results are stored in the VEM library.
[0291] It is reminded to the user of the present application that the system of the present application is beneficial to realize cloud mode + fixed terminal or mobile terminal, wherein the VEM library and VEM parameters can not only be stored in the cloud center and the terminal, but also can be stored in the nodes of the block chain in the block chain mode. Therefore, the user is managed in the membership mode, and the convenience of intellectual property management is provided for the user.
[0292] 10、Others
[0293] On the basis of the foregoing scheme, the modified model of the present application further includes but is not limited to the following extension steps or methods:
[0294] ST700: The modified model also includes application systems of mobile terminals and PC terminals, and application systems of cloud mode and block chain.
[0295] ST800: The modified model also includes a support hardware system, including a communication interface, a recording module, a tuning module, a playback module, an encryption module, a decryption module, an interface for a TikTok system, an interface for a video number, and a support AI karaoke system.
[0296] ST900: According to the application of the subsequent large model, a synchronous signal is provided to access the AI system including DeepSeek, Kimi.AI, and ChatGPT, forming an AI intelligent agent Agent.
[0297] STA00: The modified model also includes an interface protocol, providing MIDI protocol, MSC extended protocol, and OSC network protocol based on hardware and network, providing AES3 / PDIF protocol, MADI protocol based on transmission layer, and providing network audio transmission protocol such as Dante protocol, AVB / TSN protocol, and AES67 protocol.
[0298] Among them, MIDI is Musical Instrument Digital Interface, MSC is MIDI Show Control, OSC is Open Sound Control, AES3 / PDIF is a point-to-point digital audio transmission standard, MADI is Multichannel Audio Digital Interface, Dante is an automatic discovery, low latency, high channel number protocol developed by Audinate company, AVB / TSN is an audio and video low latency transmission protocol, and AES67 is an interoperability standard protocol based on existing network technology.
[0299] It is particularly pointed out that in the specification, the variable names (such as VEM-Token with a number suffix) are inconsistent before and after the steps according to the habits of computer field marking, and the same variable (such as VEM-Token1.1) changes with the change of steps, and the calculation result of the variable name will change, which should be understood by the engineering and technical personnel in the field.
Claims
1. A method for constructing a VEM-token vocal emotion multimodal magic modification model, characterized in that, Comprise: ST100: Collect sample files and user files, according to the VEM-Token model, adopt beat capture and VEM-Token segmentation, respectively obtain VEM-Token1 sequence of sample files and VEM-Token2 sequence of user files, according to VEM-Token1 sequence, adopt beat alignment for all VEM-Token2 sequences, generate aligned VEM-Token2 sequence; ST200: According to the VEM parameters included in the VEM-Token model, identify the VEM parameters of the VEM-Token1 sequence, determine the modification scheme by the user, process the VEM parameters of the aligned VEM-Token2 sequence according to the parameters of the VEM-Token1 sequence, and generate the VEM-Token2 sequence of the result file that conforms to the style of the VEM-Token1 sequence and conforms to the modification scheme.
2. The method of claim 1, wherein, Specifically comprise: ST110: For sample files not including loop segments, adopt beat capture to mark the start and end of the beat, and generate all VEM-Token1 sequences; ST120: For sample files including loop segments, from the second loop segment to the end of all loop segments, after beat capture, perform start point fine-tuning step and end point fine-tuning step to realize beat alignment within the loop segment, and generate all VEM-Token1 sequences; ST130: According to the VEM-Token1 sequence, perform beat capture and beat alignment steps including start point fine-tuning step and end point fine-tuning step in all VEM-Token2 sequences to process the corresponding beat, and generate VEM-Token2 sequence; ST140: User files specifically include: user files produced by the user imitating the sample file to sing, user files produced by completely cloning the vocal characteristics of the sample file, and result files produced by mixing locally sung user files and locally cloned user files.
3. The method of claim 2, wherein, Further specifically comprise: ST210: The VEM-Token model further comprises VEM library and VEM processor; ST211: VEM parameters include VEM classification, VEM function and established VEM coordinate system recording multiple modalities in vocal emotion; ST212: Modalities include one or a combination of lyrics, singing, accompaniment, commentary, vocal style, music, emotional basis, accompaniment instrument, video, image, and environmental sound; ST213: Emotional basis includes one or a combination of joy, sorrow, sadness, anger, fear, disgust, surprise, calm, longing, expectation, trust, love, hate, emotion, and hatred; ST214: Vocal style includes one or a combination of national song singing method, popular song singing method, rock song singing method, western song singing method, pop song singing method, original song singing method, and opera singing method; ST215: Vocal style includes the combination of fundamental frequency and overtone emitted by vocal cords, resonance cavity during singing, speaking and reciting by the user, which is the inherent characteristic of the user that distinguishes him from others. ST216: The VEM classification collects a variety of vocal files, and the vocal files are judged by human vocal experts or learning algorithms in terms of singing and accompaniment, supervised learning and deep learning are used to train the VEM function to obtain the VEM parameters, which are added to the VEM library; ST220: The VEM processor provides an operation interface for user and machine interaction to complete the specific implementation of the magic modification scheme.
4. The method of claim 3, wherein, The magic modification scheme includes the steps of singing magic modification, accompaniment sound magic modification, emotion overtone magic modification, and emotion fluctuation magic modification, specifically including: ST310: The magic modification scheme includes singing magic modification, specifically including: ST311: The VEM processor is used for preprocessing sample files and user files, setting a singing filter, converting sample files and user files into frequency spectrum format files, and separating VEM-Token1.1 sequences of singing in sample files and VEM-Token2.1 sequences of singing in user files; ST312: According to the beat capture and beat alignment model, the beat start points and beat end points of all VEM-Token1.1 sequences and all VEM-Token2.1 sequences are captured respectively, and the beat start points and beat end points of VEM-Token2.1 sequences are aligned with the beat start points and beat end points of VEM-Token1.1 sequences using a start point fine-tuning model and an end point fine-tuning model; and / or, ST320: The magic modification scheme also includes accompaniment sound magic modification, specifically including: ST321: The VEM processor is used for preprocessing sample files and user files, setting an accompaniment filter, converting sample files and user files into frequency spectrum format files, and separating VEM-Token1.2 sequences of accompaniment sound in sample files and VEM-Token2.2 sequences of accompaniment sound in user files; ST322: According to the beat capture and beat alignment model, the beat start points and beat end points of all VEM-Token1.2 sequences and all VEM-Token2.2 sequences are captured respectively, and the beat start points and beat end points of VEM-Token2.2 sequences are aligned with VEM-Token1.2 sequences using a start point fine-tuning model and an end point fine-tuning model; and / or, ST330: The magic modification scheme also includes emotion overtone magic modification, specifically including: ST331: The VEM processor is used for preprocessing sample files and user files, setting an emotion overtone filter, converting sample files and user files into frequency spectrum format files, and separating VEM-Token1.3 sequences of emotion overtones in sample files and VEM-Token2.3 sequences of emotion overtones in user files; ST332: According to the beat capture and beat alignment model, the beat start points and beat end points of all VEM-Token1.3 sequences and all VEM-Token2.3 sequences are captured respectively, and the beat start points and beat end points of VEM-Token2.3 sequences are aligned with VEM-Token1.3 sequences using a start point fine-tuning model and an end point fine-tuning model; and / or, ST340: The modification scheme further includes emotional fluctuation modification, specifically including: ST341: Using a VEM processor to preprocess the sample file and the user file, set an emotional fluctuation filter, and convert the sample file and the user file into a frequency spectrum format file, separate the VEM-Token1.4 sequence of the emotional fluctuation of the sample file and the VEM-Token2.4 sequence of the emotional fluctuation of the user file; ST342: According to the beat capture and beat alignment model, capture the beat start point and beat end point of all VEM-Token1.4 sequences and all VEM-Token2.4 sequences respectively, and use the start point fine-tuning model and the end point fine-tuning model to align the beat start point and beat end point of the VEM-Token2.4 sequence with the VEM-Token1.4 sequence.
5. The method of claim 4, wherein, The modification scheme further includes learning to sing modification steps, specifically including: ST410: Learning to sing modification includes pure learning to sing, specifically including: ST411: Using a VEM processor, learning to sing more than once by a user in units of a measure composed of multiple beats based on all or part of the sample file, and converting the multiple sets of VEM-Token2.1 sequences recorded by the user into user files, and then selecting by the user the most optimal set of VEM-Token2.1 sequences from the multiple sets of VEM-Token2.1 sequences as the selected VEM-Token2.1 sequences; ST412: Using the steps of ST310, ST311, and ST312 to generate the beat-captured and aligned VEM-Token2.1 sequence from the selected VEM-Token2.1 sequence; and / or, ST420: Learning to sing modification also includes mixed learning to sing, specifically including: ST421: Setting weights A and B for the VEM-Token1.1 sequence and the VEM-Token2.1 sequence, and using the formula A x VEM-Token1.1 sequence + B x VEM-Token2.1 sequence to calculate the VEM-Token2.1 sequence after mixed learning to sing, where A is less than 0.3 and A + B = 1.
0.
6. The method of claim 4, wherein, The modification scheme further includes vocal cloning modification steps, specifically including: ST430: The modification scheme further includes vocal cloning modification, specifically including: ST431: Querying the VEM library, if there are VEM parameters of the customer's vocal style, obtaining the VEM parameters of the customer's vocal style; or, If there are no VEM parameters of the customer's vocal style, or the existing VEM parameters of the customer's vocal style do not meet the user's requirements, collect a user's singing recording of a practice song, a speaking and reciting recording, convert the recording into a frequency spectrum format file with the user's vocal style using a VEM processor, an emotional overtone filter, and an emotional fluctuation filter, obtain the VEM parameters of the user's vocal style according to VEM classification, and store them in the VEM library; ST432: According to the VEM-Token1.1 sequence, the lyrics spectrum 1 in the sample file is recognized by using the voice recognition included in the VEM-Token model, and the lyrics spectrum 1 includes the lyrics and the start and end positions of the lyrics in the beat; ST433: The lyrics spectrum 1 of the sample file is copied as the lyrics spectrum 2 of the user file, and the voice cloning magic modification of the user is completed word by word according to the VEM parameters of the user's voice style, the lyrics spectrum 2 of the user file and the start and end positions of the beat by using the voice synthesizer, becoming the VEM-Token2.1 sequence.
7. The method of claim 6, wherein, The magic modification scheme also includes a lyrics magic modification step, which specifically includes: ST500: When the lyrics spectrum 2 is inconsistent with the lyrics spectrum 1 or the user needs to modify it, the lyrics magic modification is performed, which specifically includes: ST510: The lyrics spectrum 2 and the lyrics spectrum 1 are decomposed into lyrics sentences 2 and lyrics sentences 1 according to the semantic syntax, and the following steps are performed: ST511: The number of words in the lyrics sentences 2 and the lyrics sentences 1 is the same, the lyrics spectrum 2 is copied as the lyrics spectrum 1 according to the beat, and each word in the lyrics in the lyrics sentences 2 is filled in the corresponding beat position one by one, or manually modified by the user to become the modified lyrics spectrum 2; ST512: The number of words in the lyrics sentences 2 and the lyrics sentences 1 is different, and the lyrics of the lyrics 2 are manually modified and filled in the corresponding beat position one by one to become the modified lyrics spectrum 2; ST513: The voice cloning magic modification of the user is completed word by word according to the VEM parameters of the user's voice style, the lyrics spectrum 2 and the start and end positions of the beat by using the voice synthesizer, becoming the VEM-Token2.1 sequence.
8. The method according to claim 5 or 7, characterized in that, The magic modification scheme also includes the steps of pitch calibration magic modification, ornamentation magic modification, beat length magic modification, rhythm speed magic modification, beat strength magic modification, timbre magic modification, emotion magic modification, video magic modification, free magic modification, multiple sample magic modification, real-time listening of what you hear, which specifically include: ST520: Pitch calibration magic modification: According to the musical theory of twelve equal temperament, the fundamental frequency of the sound must be equal to the node frequency of the twelve equal temperament, and the frequency of the sound between the two adjacent node frequencies needs to be adjusted up or down to the node frequency; ST530: Front ornamentation magic modification: In a beat, when the beat start point of VEM-Token2.1 and the corresponding beat start point of VEM-Token1.1 fall within 1 / 2 beat on the time axis, the ornamentation is used to compensate before the beat start point of VEM-Token2.1 to align the start point, and the ornamentation includes tremolo, slide, extension, and breathing sound; ST540: Post-ornamentation magic modification: In a beat, when the beat end point of VEM-Token2.1 and the corresponding beat end point of VEM-Token1.1 lead by 1 / 2 beat on the time axis, the ornamentation or rest is used to compensate after the beat end point of VEM-Token2.1 to align the end point; ST550: Beat length modification: when the length of the beat of VEM-Token2.1 does not match the length of the beat of the corresponding VEM-Token1.1, a time stretching algorithm step or a decoration note modification step is used to compress or extend the beat of VEM-Token2.1 to align with the beat of the corresponding VEM-Token1.1; ST560: Tempo fast / slow modification: when the overall tempo of the user file needs to be sped up or slowed down, a time stretching algorithm step is used to synchronously compress or extend the tempo of the vocals and the accompaniment of the user file; ST570: Beat strong / weak modification: for VEM-Token2.1 and / or VEM-Token2.2, different processing based on the base layer and the emotion layer is used according to the user's needs, wherein: ST571: The base layer includes a step of volume dynamic processing and envelope shaping, which adjusts the threshold for VEM-Token2.1 and VEM-Token2.2; ST572: The emotion layer includes intelligent dynamic control based on AI / machine learning, which specifically includes training a model to intelligently identify the beat, instruments in the audio, and automatically generating dynamic processing parameters according to pre-set emotion labels or target loudness curves, and the training results are stored in the VEM library; ST580: When the length of the beat of VEM-Token2.1 changes, the length of the beat of VEM-Token2.2 needs to be synchronously verified, and when the lengths of the beats are inconsistent, time stretching is used to capture and align VEM-Token2.1 and VEM-Token2.2; ST590: Tone modification, specifically including: ST591: for VEM-Token1.1 sequence and VEM-Token2.1 sequence, a filter including multiple groups of high-order harmonics of fundamental frequency is adopted, and the frequencies of the multiple groups of high-order harmonics are decomposed: F1, F2, F3, …, F n , the amplitudes of the high-order harmonics are decomposed: A1, A2, A3, …, A n , where n is the order of the harmonic, and n is less than 50; ST592: using the steps of volume dynamic processing and envelope shaping, respectively amplifying or reducing A1, A2, A3,..., A n the amplitude of the overtones of one or more specified harmonic components to change the timbre of the user file; or, ST593: Querying AI / machine learning tone dynamic processing parameters in the VEM library, adjusting the dynamic processing parameters to change the tone of the user file; ST5A0: Emotion modification, specifically including: based on the user's emotion modification needs, querying the VEM library with VEM-Token2.1 sequence as the independent variable, adjusting the VEM parameters including emotion modification, vocal style, and vocal style needs, to obtain the emotion modification result; ST5B0: Video modification, when the user file needs to be adapted to the video, the content and tempo of the video are adjusted according to the VEM parameters and the tempo to adapt to the needs of the user file; ST5C0: Free modification, the user modifies the content and tempo of VEM-Token2.1, VEM-Token2.2, and the video according to the VEM parameters and one or more modalities to adapt to the needs of the user file; ST5D0: Multiple sample modification, the user selects one or more sample files and selects part of the parameters in the VEM parameters corresponding to part of the sample files and selects another part of the parameters corresponding to another part of the sample files, and modifies the content and tempo of VEM-Token2.1, VEM-Token2.2, and the video to adapt to the needs of the user file; ST5E0: Real-time monitoring of what is heard, including real-time monitoring, evaluation, and scoring of the modified results by the user, and submitting the VEM-Token2 sequence of the modification process and the modified results to the VEM library.
9. The method of claim 8, wherein, The modification model also includes member management, including: ST600: Apply for membership in the modification model according to user needs, establish member profiles, and store them in the VEM library. ST610: Member profiles include user information, sample file information, user file information, VEM parameters, voiceprint encryption, and voiceprint decryption, where the encryption and decryption keys include member signatures, member images, member videos, and member VEM parameters. ST620: Member management includes forward, backward, rollback, add, delete, query, modify, store, and maintain operation steps during the modification process. ST630: Member management also includes real-time modification, real-time monitoring, real-time scoring, supervised learning, reinforcement learning, rewards, and punishments of user files, with results stored in the VEM library.
10. The method of claim 9, wherein, The modification model also includes: ST700: The modification model also includes mobile application systems and PC application systems, as well as cloud-based application systems and blockchain application systems. ST800: The modification model also includes supporting hardware systems, including communication interfaces, recording modules, tuning modules, playback modules, encryption modules, and decryption modules, as well as interfaces for Douyin systems, interfaces for WeChat videos, and support for AI karaoke systems. ST900: Based on providing synchronization signals to subsequent large model applications, AI systems including DeepSeek, Kimi.AI, and ChatGPT are accessed to form AI agents. STA00: The modification model also includes interface protocols, providing MIDI protocols, MSC extension protocols, and OSC network protocols based on hardware and networks, providing AES3 / PDIF protocols and MADI protocols based on transmission layers, and providing network audio transmission protocols such as Dante protocols, AVB / TSN protocols, and AES67 protocols.
Citation Information
Patent Citations
VEM-Token beat capture and alignment model construction method
CN120748450A
Cited By
VEM-Token world model robot expression function construction method
CN120951102A
Method for constructing robot expression function of vem-token world model
CN120951102B