Method for hierarchical fusion of vem-token emotion synchronization functions
By employing a layered fusion method based on the VEM-Token emotion synchronization function, the shortcomings of NLP-token in parsing non-textual information modal emotions are addressed, enabling efficient and dynamic music emotion recognition and understanding, and supporting the application of intelligent agents.
Patent Information
- Application Number
- CN202511428558.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-01
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-10-01
AI Technical Summary
Existing NLP-token-based technologies cannot accurately parse and understand the emotions of non-textual information modalities such as singing, speech, and recitation in real time. This results in inaccurate emotion classification boundaries, difficulty in describing spatial correlations, and difficulty in capturing the continuity of emotion changes over time, creating an illusion of large model outputs.
The VEM-Token emotion synchronization function hierarchical fusion method is adopted. Through the VEM-Token vocal emotion multimodal model, music beats are used as information words. Combined with multi-layer weighted scanning, recurrent neural network, long short-term memory network and self-attention mechanism, the emotion synchronization function of song file is calculated to realize the automatic recognition and management of emotions.
It implements function value mapping of high-dimensional emotion names, emotion values, and emotion weights, transcends the discrete description of NLP-Tokens, dynamically captures musical emotions, and provides a guarantee for generating text, emotion-generated text, text-generated animation, and emotion-generated actions, supporting the application of intelligent agents.
Smart Images

Figure CN120913602B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to the model construction and processing of AI music, singing voice, speech, recitation, disease patient voice change and AI intelligent agent Agent, and more particularly to an innovative emotion synchronization function design and operation method. Unlike the traditional NLP-token word element system, the present application depends on the Internet text for the analysis of music, singing voice and speech emotion. The present application first proposes a new system based on VEM-Token vocal music voice emotion recognition, adopts a VEM-Token emotion synchronization function hierarchical fusion method, and reconstructs the analysis model system. BACKGROUND
[0002] In the field of artificial intelligence, so far the decomposition of information is still based on the natural language word element division method of NLP-Token (Natural-Language-Processing Token, natural language processing word element). If the analysis and application are based on text information modalities, NLP-Token has natural advantages. Large models have learned all the books and web pages based on human text and memorized them in large models. However, if the information modalities of non-text, such as singing voice, speech, recitation emotion, music style, etc., today's large models still search for text descriptions learned from books and web pages in the past in the model memory. Through these text descriptions, "explanation" and "understanding" are obtained, that is, the information modalities based on NLP-Token word elements are still used. Since we cannot know where the large model obtains the corpus and cannot predict the correctness of the corpus, the "illusion" of the large model cannot be avoided. In addition, many information modalities in nature cannot be accurately described directly through language and text. For example, human emotions and even animal emotions, according to the classification of psychology, there are at least joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, emotion, and hatred. By carefully examining these emotions, we will find that:
[0003] (1) The classification boundary of emotion cannot be accurately described by text. For example, it is difficult to accurately define the classification boundary between "joy" and "love", and how to define this ambiguity and convertibility in text description?
[0004] (2) There is a spatial correlation between emotions, which is difficult to describe in words. Even if we accept the above psychological classification, for example, "fear", "disgust" and "surprise" are different categories, but how do they affect each other and how are they related? Although NLP-Token can measure the angle and length of the vector in the high-dimensional coordinate system based on the dimension of the emotion category, what if the coordinate system itself is flawed?
[0005] (3) Emotions are continuous analog functions that change over time, and written descriptions can only be a slice of this function and cannot be described in real time. For example, when a song describing "sadness" and "hate" is performed, the singer's emotions will fluctuate as the rhythm flows, and different singers, different environments, and even different performances will have different fluctuations. For example, Xianglin'ao repeatedly tells the story of Mao being eaten by a wolf, sometimes with tears of sadness, and sometimes with tears of sadness.
[0006] By analogy, we can easily find that even if the big model learns more knowledge, it cannot track emotions in real time and output accurate "explanations" and "understandings". The root cause of this situation is the NLP-Token model itself.
[0007] The team of the present inventors first proposed a brand-new model design, which includes a piece of granted Chinese invention patent "VEM-Token vocal emotion multimodal tokenization singing and accompaniment deep learning method, CN120126506" (hereinafter referred to as "VEM-Token vocal emotion multimodal model") and a piece of invention patent under review "Method for constructing VEM-Token rhythm capture and alignment model, 202511249168.0" (hereinafter referred to as "VEM-Token rhythm model"), which successfully segments music files by taking music rhythm as information token, i.e. VEM-Token (Vocal-Emotion-Multimodal Token), which is different from the traditional NLP-token method, but an innovative emotion synchronization function VEM-sync, which synchronizes with the VEM-Token sequence of the music rhythm, directly describes high-dimensional multiple emotions in mathematical language, thereby avoiding the deviation and illusion of natural language description of high-dimensional emotion functions. The present invention includes VEM-sync and synchronous content definition, rhythm attribute, emotion attribute, emotion function definition, one or combination of multi-layer weighted scanning, recurrent neural network, long short-term memory network, self-attention mechanism, retrieval enhancement generation RAG network, hierarchical fusion calculation outputs the time emotion function of the vocal file, according to this sound function or emotion function, directly drives the sound text / emotion text, sound expression / emotion expression, sound multi-dimensional animation / emotion multi-dimensional animation, supports model context protocol MCP and function calling Function Calling, provides support for copyright management and encryption / decryption or interface.
[0008] The present invention application is a patent pool patent of CN120126506 and 202511249168.0, which can be connected to an AI system or an independently developed application system, and developed into a vocal intelligent agent Agent that can listen to music and recognize music scores. The model obtains the VEM library through supervised learning. Due to the existence of this professional and accurate VEM library learned by human experts, the possibility of "illusion" produced by the large model in the application almost disappears.
[0009] In addition, as a VEM-Token-based patent pool patent, the present invention has expansion potential in supporting embodied robots facial expressions, supporting virtual digital humans, and supporting digital user self.
[0010] The VEM-Token concept mentioned in this application is based on the basic concept and steps defined in CN120126506, 202511249168.0 invention patent, unless otherwise emphasized or specifically defined in this application. Among them, "vocal music file", "music file", "song", "music" mentioned in this application file have the same meaning unless otherwise stated.
[0011] The present application aims to apply to the current research and application of large models, such as OpenAI, DeepSeek, Google Gemini, Kimi, bean bag large model, Wenxin Yanyan, etc. After connecting the front end and some applications of the back end, various intelligent agents are formed. The present application aims to access these large models, realize bidirectional communication with them, form music / vocal music-based artificial intelligence applications, and even music AI agents to broaden the application of AI and provide powerful innovation and support.
[0012] Disadvantages of prior art methods
[0013] (1) The existing NLP-token-based technology analyzes transitions and describes emotions through text. Text is discrete and has clear boundaries, while emotions are not only continuous analog quantities and fuzzy concepts, but also time functions that change over time. Therefore, using NLP-token to describe emotions is not a good approach, as it will lose a lot of information details in emotions.
[0014] (2) Currently, there is no mathematical model of voice / music emotion generation function, i.e., "sound function / emotion function" in common language. Therefore, there is no method to generate "sound text / emotion text", "sound animation / emotion animation", and "sound action / emotion action" driven by voice / music. SUMMARY
[0015] In view of the deficiencies of the prior art, the present application proposes a novel VEM-Token emotion synchronization function hierarchical fusion method. This method is based on the VEM-Token vocal music emotion multi-modal model, rather than the NLP-token model, to automatically identify, manage, and process vocal music emotions, achieving the purpose and intent of the present application and providing an effective preliminary basis for realizing "sound text / emotion text", "sound animation / emotion animation", and "sound action / emotion action".
[0016] The purpose and intent of the present application are achieved by using the following technical solutions and working steps:
[0017] 1. VEM-Token emotion synchronization function hierarchical fusion method
[0018] The present application is a method for hierarchical fusion of VEM-Token emotional synchronization functions, including but not limited to the following steps:
[0019] ST100: Adopting a VEM-Token vocal emotion multi-modal model, dividing the song file into VEM-Token sequences in units of beats, setting the emotional synchronization function as VEM-sync, wherein VEM-sync includes but is not limited to a VEM-sync vector sequence corresponding to and aligned with the VEM-Token sequence, and the VEM-sync vector includes but is not limited to synchronization content and a synchronization pointer pointing to the corresponding VEM-Token beat.
[0020] ST200: According to the VEM-Token sequence, the synchronization content of the VEM-sync vector sequence is calculated and obtained, and the synchronization content includes but is not limited to beat attributes and emotional attributes, wherein the emotional attributes include but are not limited to one or more emotional names, emotional values, and emotional weights.
[0021] ST300: The emotional weight function includes but is not limited to: adopting hierarchical processing of the song file to obtain the corresponding VEM-sync vector sequence, and the emotional values therein include but are not limited to hierarchical emotional weights, and adopting one or a combination of multi-layer weighted scanning, recurrent neural networks, long short-term memory networks, self-attention mechanisms, and retrieval enhancement generation for forward propagation, backward propagation, and omnidirectional propagation to obtain the emotional synchronization function of the song file.
[0022] 2. VEM-Token vocal emotion multi-modal model
[0023] On the basis of the foregoing basic scheme, the present application in the VEM-Token vocal emotion multi-modal model includes but is not limited to one or more combinations of the following steps or methods of beat capture and alignment:
[0024] ST110: The time measurement of the VEM-Token sequence is in seconds or beats, and the length starts from 0 to the end of the song.
[0025] ST120: The synchronization pointer includes: a single synchronization pointer pointing to and aligned with the VEM-Token beat start time, or a double synchronization pointer pointing to the VEM-Token beat start and end times, respectively.
[0026] ST130: The length of the data structure of the synchronization content includes but is not limited to fixed length and variable length.
[0027] ST140: The VEM classification refers to the classification of emotions, including but not limited to independent emotions, opposite emotion pairs, and related opposite emotion groups, and one or a combination of one-dimensional single-direction, one-dimensional double-direction, and multi-dimensional double-direction is established in turn.
[0028] ST150: The VEM mode includes but is not limited to one or a combination of lyrics, singing, accompaniment, style, music, accompaniment instruments, video, and image.
[0029] ST160: The style includes but is not limited to one or a combination of national song singing, popular song singing, western song singing, pop song singing, original song singing, and opera singing.
[0030] ST170: The emotion name is one or a combination of joy, sadness, anger, fear, disgust, surprise, calmness, expectation, trust, love, hatred, emotion, and hatred included in the VEM mode, and the emotion value is set to the emotion measurement including but not limited to percentage or numerical value.
[0031] ST180: The VEM library includes but is not limited to VEM mode, style, emotion name, and emotion value, which are respectively generated by human vocal music experts for typical song guidance supervised learning, reinforcement learning, and VEM-sync vector sequence generated by deep learning of VEM-Token vocal emotion multi-modal model.
[0032] 3. Basic step of synchronizing content
[0033] On the basis of the foregoing scheme, the present application includes but is not limited to one or more of the following steps or methods in the basic step of synchronizing content:
[0034] ST210: The beat attribute is calculated and obtained by the VEM-Token vocal emotion multi-modal model, or is collected and obtained from the score of the song file. The beat attribute specifically includes but is not limited to tempo, measure, beat type, and type, wherein the measure is composed of more than one beat, the time unit uses the number of beats per minute, the beat type includes but is not limited to strong, secondary strong, weak, and rest, and the type includes but is not limited to 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, 3 / 8 beat, 6 / 8 beat, and 7 / 8 beat.
[0035] ST220: The emotion attribute is calculated and obtained by the VEM-Token vocal emotion multi-modal model for the VEM-Token sequence and one or more emotion names, and the emotion value of the emotion name is normalized to between -100% and +100%. The emotion weight is set, and the tempo, emotion name, emotion value, and emotion weight constitute a storage unit of an emotion weight function and are stored in the VEM-sync vector sequence.
[0036] ST230: The emotion weight function is defined as a weighted function for an emotion value to modify the resulting emotion value of an emotion name in a VEM-sync vector in the VEM-sync vector sequence in different layers.
[0037] ST240: The layers are set as the VEM-sync vector sequence for the song file, the division includes the beat layer, the measure layer, the sentence layer and the whole song layer, and the emotion weight function is applied layer by layer to obtain the corresponding weighted emotion value.
[0038] ST250: According to the emotion change of the whole song file, the emotion weight function is used to calculate the emotion synchronization function of the whole song file according to the hierarchical fusion.
[0039] 4, the multi-layer weighted scanning step of the emotion weight function
[0040] On the basis of the foregoing scheme, the multi-layer weighted scanning step of the emotion weight function includes but is not limited to one or more combinations of the following steps or methods:
[0041] ST410: The emotion attribute of the beat layer of the song file is obtained in units of beats.
[0042] ST420: The measure layer is added in the VEM-sync vector sequence in units of measures, the emotion values of the same emotion name in the measure layer are averaged, and each kind of emotion and the corresponding emotion average value are stored in the measure layer.
[0043] ST430: The sentence layer is added in the VEM-sync vector sequence in units of sentences, the emotion values of the same emotion name in the sentence layer are accumulated, the style weight is extracted from the VEM library according to the style, or the style weight is selected by the user, the weighted emotion value of the emotion name is obtained by using the weighted algorithm, and each kind of emotion and the corresponding weighted emotion average value are stored in the sentence layer.
[0044] ST440: The whole song layer is added in the VEM-sync vector sequence in units of whole songs, the emotion values of the emotion name in the sentence layer are sorted from large to small, the first part of the emotion name or the user-specified emotion name is selected as the high-weight emotion, wherein the first part is less than 40% according to the statistical result in the VEM library or selected by the user, the high-weight emotion and the corresponding emotion value are stored in the whole song layer, output as the emotion weight function, and stored in the VEM library.
[0045] 5, the step of the recurrent neural network of the emotion weight function
[0046] On the basis of the foregoing scheme, the steps of the emotion weight function of the recurrent neural network include, but are not limited to, one or more combinations of the following steps or methods:
[0047] ST510: A recurrent neural network RNN is established to realize a memory network model, realize forward propagation of emotion, and input emotion values of a VEM-sync vector of a previous beat as hidden history state of a current beat when calculating the VEM-sync vector of the current beat, combine current emotion values to perform fusion calculation, so as to enhance emotion representation of the previous and current beats and obtain current emotion values.
[0048] ST520: Emotion names and emotion values of the VEM-sync vector are obtained, a VEM-sync vector with a beat number t is marked as hidden history state VEM-sync (t) of a current beat, a VEM-sync vector with a beat number t-1 is marked as hidden history state VEM-sync (t-1) of a previous beat, and t starts from 2 until the end of a song.
[0049] ST530: A calculation formula of the memory network model is set as:
[0050] VEM-SYNC (t) = f ( W_x ▪ VEM-sync (t) + W_h ▪ VEM-sync (t-1) + b ),
[0051] wherein VEM-SYNC (t) is hidden state, f is an activation function, W_x and W_h are emotion weight matrices that can be trained, b is a bias vector, VEM-SYNC (t) output is used as an emotion weight function, and is stored in a VEM library.
[0052] 6. Steps of the emotion weight function of the long short-term memory network
[0053] On the basis of the foregoing scheme, the steps of the emotion weight function of the long short-term memory network include, but are not limited to, one or more combinations of the following steps or methods:
[0054] ST610: A long short-term memory network LSTM is established to realize a memory network model that is longer than a previous beat, realize forward propagation of emotion, and input emotion values of a VEM-sync vector of a forgotten beat or forgotten emotion as hidden history state of a current beat when calculating the VEM-sync vector of the current beat, combine current emotion values to perform fusion calculation, so as to enhance emotion representation of the previous and current beats and obtain current emotion values.
[0055] ST620: The forgotten beat is the beat that is not important, including but not limited to more than one local accompaniment overpass, accompaniment, breath in the song file, or the beat selected by the user, and the forgotten emotion is the emotion name and its emotion value that is not important, or the emotion name and its emotion value selected by the user.
[0056] ST630: The output of the memory network model is an emotion weight function, which is stored in the VEM library.
[0057] 7. Self-attention mechanism step of emotion weight function
[0058] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following model steps or methods in the self-attention mechanism step of the emotion weight function:
[0059] ST710: Obtain the beat number, emotion name and emotion value of the VEM-sync vector of the song file, and mark the VEM-sync vector with beat number t as VEM-sync(t).
[0060] ST720: Establish a Transformer encoder to realize forward propagation, backward propagation and omnidirectional propagation of emotion in the VEM-sync vector sequence of all beats of the song file through self-attention mechanism, calculate the global dependency relationship between all beat pairs in the VEM-sync vector sequence, and output a group of emotion value sequences enhanced by global up and down VEM-Token sequence weights.
[0061] ST730: The Transformer encoder includes but is not limited to more than one encoder layer, and each encoder layer further includes but is not limited to a multi-head self-attention layer and a feedforward neural network layer, which complete the update and output of the emotion weight function for all VEM-sync vector sequences through residual connection, layer normalization, input embedding and position encoding, wherein the position encoding is obtained from the beat number.
[0062] ST740: Input all VEM-sync vector sequences of the song file to the Transformer encoder at one time, and set one or more emotion names and emotion values that need to be focused on to access the input end of the multi-head self-attention layer, add a multi-head self-attention layer and a feedforward neural network layer in the VEM-sync vector, and store the calculation result of the Transformer encoder in the multi-head self-attention layer.
[0063] ST750: Set a dynamic mechanism selection to access the emotion name and emotion value of the multi-head self-attention layer.
[0064] ST760: Map the emotion name and emotion value of each beat to the Query, Key, Value spaces through a learnable weight matrix. Calculate the dot product of Query and Key on the beat sequence to obtain the association weight between beats. Weighted sum all the Value using the association weight to obtain the song file emotion weight information.
[0065] ST770: For emotion names and emotion values not included in the multi-head self-attention layer input, set the prompt word and use the RAG step to generate, select the top N emotion names, emotion values and beat numbers with the largest emotion values, and incorporate them into the emotion weight function to prevent omission, where the value of N is determined according to the song file or by the user.
[0066] ST780: Output the result stored in the multi-head self-attention layer as the emotion weight function and store it in the VEM library.
[0067] 8, Enhanced RAG step
[0068] On the basis of the foregoing scheme, the present application further comprises one or more combinations of the following steps or methods in the output step of the emotion weight function:
[0069] ST810: Obtain the beat number, emotion name and emotion value of the VEM-sync vector sequence of the song file, and mark the VEM-sync vector with beat number t as the VEM-sync(t) array, where the value of t includes one or more paragraphs of the song file.
[0070] ST820: Retrieve the VEM library according to the VEM-sync(t) array to obtain one or more VEM-sync vector sequences with K-RAG similar emotion values as the reference VEM-sync(t) array sequence, and the similarity is determined by the user.
[0071] ST830: Take the VEM-sync(t) array and the reference VEM-sync(t) array as the input of the multi-head self-attention mechanism of the Transformer encoder, and use the self-attention mechanism to calculate the cross-attention weight between the VEM-sync(t) array and the reference VEM-sync(t) array, and then perform weighted fusion to obtain the fusion result VEM-sync(t) array.
[0072] ST840: Use the fusion result VEM-sync(t) array to verify or correct the preliminary emotion analysis result calculated by the multi-head self-attention network.
[0073] 9. The step of outputting the mood weight function
[0074] Based on the foregoing scheme, the present application further comprises, but is not limited to, one or more combinations of the following steps or methods in the step of outputting the mood weight function:
[0075] ST910: Converting the mood score output in written format, digital format or MIDI communication interface format according to the mood weight function of the entire VEM-sync vector sequence of the song file.
[0076] ST920: Establishing a format for communication with the Agent end, and outputting the mood score to the Agent end according to the mood weight function of the entire VEM-sync vector sequence of the song file.
[0077] ST930: Establishing an NLP-token-oriented interface according to the standard of NLP-token words in the natural language model, including a prompt word interface, a context protocol MCP interface and an AI interface.
[0078] ST950: For a song file with video and background pictures that are synchronized in time, establishing a bidirectional picture pointer one that points to and aligns the beat start point and the picture synchronization point, respectively, according to the tempo and beat start point of the VEM-Token sequence of the song file.
[0079] ST960: Adding a picture pointer two in the VEM-sync vector sequence, which points to and aligns the picture synchronization point according to the tempo and beat start point of the VEM-Token sequence.
[0080] ST970: Establishing a synchronization and alignment relationship between the VEM-sync vector sequence and the video and background pictures through the tempo and beat start point of the VEM-Token sequence, the synchronization pointer, the picture pointer one and the picture pointer two.
[0081] ST980: Outputting fusion information from the VEM-sync vector sequence to the video and background pictures according to user requirements, including but not limited to lyrics, song scores and information or actions in the VEM-sync vector sequence that meet user requirements.
[0082] ST990: Outputting fusion information from the VEM-sync vector sequence to 3D animation expression actions, language actions and body actions, including but not limited to, according to user requirements, to drive 3D animation and expression actions, language actions and body actions to make expressions and body actions corresponding to the mood vectors in the VEM-sync vector sequence.
[0083] ST9A0: support model context protocol MCP, including MCP Host, MCP Client, MCP Server interface, provide function calling function.
[0084] 10. Intellectual property record step
[0085] On the basis of the foregoing scheme, the application in the step of video and background picture fusion, specifically includes the following one or more combinations of steps or methods:
[0086] STA10: insert the mark vector space CV or interface of intellectual property record in the whole VEM-sync vector sequence of song file, for storing copyright vector.
[0087] STA20: copyright vector includes but is not limited to copyright holder information, version number, authorization information, encryption and decryption mode.
[0088] STA30: including encryption and decryption engine, including but not limited to elliptic algorithm, public key, private key mode.
[0089] STA40: managed by using blockchain and cloud mode.
[0090] 11. Invention purpose and intention
[0091] The purpose and intention of the VEM-Token emotion synchronization function hierarchical fusion method of the application are:
[0092] (1) realize the mapping of the VEM-Token sequence of the song file into the VEM-sync emotion synchronization function, and obtain the function values of the high-dimensional emotion name, emotion value and emotion weight through hierarchical and fusion.
[0093] (2) cross the transition of discrete text description of NLP-Token, thereby avoiding the deviation of the discrete natural language text for high-dimensional emotion analog quantity function description.
[0094] (3) realize dynamic and continuous music emotion, provide guarantee for generating "sound function / emotion function", "sound text / emotion text", "sound animation / emotion animation" and "sound action / emotion action", and realize the detailed identification and understanding based on vocal music emotion.
[0095] (4) create the automation and intelligence of music file emotion identification and understanding, realize the multi-modal quantitative representation ability of vocal music emotion.
[0096] (5) Innovate a music / vocal music-oriented model, access large models to form music AI application agents, or dedicated application systems, to provide strong support for expanding AI applications.
[0097] 12. Advantages of the invention
[0098] (1) The VEM-sync emotion synchronization function is used to realize the function value of the high-dimensional emotion name, emotion value, and emotion weight in the VEM-Token sequence describing the song file.
[0099] (2) Avoids the deviation of discrete natural language text in describing high-dimensional emotion simulation function.
[0100] (3) Realizes dynamic and continuous music emotions, provides support for generating "sound function", "sound text", and "sound animation", and realizes detailed identification and understanding based on vocal music and music emotions. BRIEF DESCRIPTION OF DRAWINGS
[0101] LIST OF DRAWINGS
[0102] Figure 1 : Emotion synchronization function diagram
[0103] Figure 2 : Emotion vector diagram
[0104] Figure 3 : Multi-layer scanning synchronization content diagram
[0105] Figure 4 : Self-attention mechanism and enhanced retrieval generation model diagram
[0106] Figure 5 : Emotion weight synchronization diagram
[0107] Figure 6 : Full song emotion synchronization function fusion diagram
[0108] DETAILED DESCRIPTION OF DRAWINGS
[0109] The detailed description of each drawing is described in the corresponding drawing number analysis section in the following
specific implementation
[0110] The present application is a patent pool patent of a granted Chinese invention patent "VEM-Token vocal emotion multi-modal tokenization singing and accompaniment deep learning method, CN120126506" and a published invention patent application "VEM-Token beat capture and alignment model construction method, 202511249168.0", and the focus is to invent a new method and a new AI model of hierarchical fusion of VEM-Token emotion synchronization function VEM-sync based on synchronization and alignment to the VEM-Token model of the music file.
[0111] The purpose and intention of the present application can be achieved by the following specific embodiments. It needs to be particularly pointed out that each specific embodiment has specific purposes and industrial applicability. Therefore, the following embodiments cannot include all the features and steps of the present application, nor constitute a limitation on the present application. The description of the claims of the present application is the summary of the invention.
[0112] The specific embodiments of the present application are as follows:
[0113] The innovative hierarchical fusion method of VEM-Token emotion synchronization function VEM-sync is a high-precision non-textual music recognition system.
[0114] Diagram explanation
[0115] The contents of the present embodiment mainly include but are not limited to the following main schematic drawings, which are: Figures 1 to 6 .
[0116] Implementation step explanation
[0117] The method steps of the present embodiment mainly include steps 1 to 10. Each of the 10 parts includes several sub-steps. Unless otherwise specified, these sub-steps are not completely required, and unless otherwise specified, the order is not essential, but according to the needs of some specific tasks, the patent implementer makes optimized and further choices.
[0118] The specific work steps are as follows:
[0119] 1. Basic scheme
[0120] The present application as a hierarchical fusion method of VEM-Token emotion synchronization function includes but is not limited to the following steps:
[0121] ST100: adopt VEM-Token vocal emotion multi-modal model, divide song file into VEM-Token sequence in time unit, set emotion synchronization function as VEM-sync, wherein VEM-sync includes but is not limited to VEM-sync vector sequence corresponding to and aligned with VEM-Token sequence, VEM-sync vector includes but is not limited to synchronization content and synchronization pointer pointing to corresponding VEM-Token beat.
[0122] ST200: according to VEM-Token sequence, calculate synchronization content of VEM-sync vector sequence, synchronization content includes but is not limited to beat attribute and emotion attribute, wherein emotion attribute includes but is not limited to one or more including but not limited to emotion name, emotion value, emotion weight.
[0123] ST300: emotion weight function includes but is not limited to: adopt hierarchical processing of song file to obtain corresponding VEM-sync vector sequence, emotion value thereof includes but is not limited to hierarchical emotion weight, adopt including but not limited to multi-layer weighted scanning, recurrent neural network, long short-term memory network and self-attention mechanism for forward propagation, backward propagation and omnidirectional propagation, and fuse to obtain emotion synchronization function of song file.
[0124] Among them, the VEM-Token vocal emotion multi-modal model refers to the model in "VEM-Token vocal emotion multi-modal tokenization song and accompaniment deep learning method, CN120126506", which specifically includes steps 1-4:
[0125] (1) adopt one or more modalities to record emotion, mark vocal emotion multi-modal as VEM, construct VEM classification, VEM coordinate system, VEM function and VEM library, vocal emotion includes one or combination of joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, emotion and hatred, multi-modal includes one or combination of lyrics, song, accompaniment, vocal style, music, emotion basis, accompaniment instrument, video and image, VEM coordinate system includes coordinate axis system established according to independent emotion, opposite emotion pair and associated opposite emotion group.
[0126] (2) collect vocal samples according to VEM classification, judge emotion of vocal samples on song and accompaniment by human vocal experts, train VEM function to obtain VEM parameters by adopting supervised learning and deep learning, and add to VEM library.
[0127] (3) Adopting VEM processor to beat mark vocal files, and separate song stream and accompaniment stream, according to beat, VEM-Token segmentation is carried out on vocal files, song stream is converted into VEM-Token1 sequence, and accompaniment stream is converted into VEM-Token2 sequence and added to the preprocessing library.
[0128] (4) Adopting deep learning, respectively generating lyrics spectrum, VEM-Token song spectrum, VEM-Token accompaniment spectrum and VEM-Token music score.
[0129] Among them, the model of VEM-Token beat capture and beat alignment refers to the model in "VEM-Token beat capture and alignment model construction method, 202511249168.0", which specifically includes steps 5-7:
[0130] (5) For vocal files, according to the VEM-Token vocal emotion multi-modal model, set the beat model including beat capture and beat alignment, capture the beat of the vocal file, and divide the vocal file into VEM-Token sequence according to the beat, and mark the position of the start of the beat and the end of the beat in each VEM-Token.
[0131] (6) Set the start point alignment model, including:
[0132] Divide the sample files included in the vocal files and the user files produced by the user imitation sample files into VEM-Token1 sequence and VEM-Token2 sequence, and according to the start point of each VEM-Token1, adopt the start point fine-tuning step to adjust the start point of VEM-Token2 at the corresponding position one by one, so that the start point of VEM-Token1 at the corresponding position is aligned.
[0133] For each segment of the cycle segment included in the vocal file, take the first segment as a reference, starting from the second segment, adopt the start point fine-tuning step to adjust the start point of each VEM-Token of each segment one by one, so that the start point of VEM-Token at the corresponding position of the first segment is aligned, until all the cycle segments end.
[0134] (7) Set the end point alignment model, including:
[0135] According to the end point of each VEM-Token1, adopt the end point fine-tuning step to adjust the end point of VEM-Token2 at the corresponding position one by one, so that the end point of VEM-Token1 at the corresponding position is aligned.
[0136] For each segment of the loop segment, taking the first segment as a reference, starting from the second segment, the end point fine adjustment step is adopted to adjust the end point of each VEM-Token of each segment one by one, so that the end points of the VEM-Tokens at the corresponding positions of the first segment are aligned, until the end of all loop segments.
[0137] In this application, although the construction of the emotion synchronization function or model is inspired by the ideas and inspirations of the two patents, it is also an independent invention. That is, after other structured models of song files have realized the emotion summary, further realization of structured, ordered and post-structured analysis of such structured systems is realized.
[0138] As Figure 1 This is a schematic diagram of the emotion synchronization function of the present application, which is analyzed as follows:
[0139] Figure 1 The lower part is the VEM-Token sequence divided by the song file, and the file attribute is the spectrum format file. The upper part is the emotion synchronization function VEM-sync, and the beat is the beat. N is the number of the sequence. The design intention is:
[0140] (1) VEM-Token is a division of song files in a certain structured, including VEM-Token vocal emotion multi-modal tokenization song and accompaniment deep learning method, CN120126506, VEM-Token beat capture and alignment model construction method, 202511249168.0 VEM-Token vocal emotion multi-modal division model, and other division methods such as MIDI, MCS model. On the VEM-Token sequence of the song file, a VEM-sync vector sequence is attached, wherein the VEM-sync vector sequence corresponds to the VEM-Token sequence one by one.
[0141] (2) The correspondence between the VEM-sync vector sequence and the VEM-Token sequence is constrained by the synchronization pointer. It should be noted here that although Figure 1 the VEM-sync vector sequence and the VEM-Token sequence in
[0142] (3) Since the VEM-Token sequence is divided according to the music beat, and the beat is the most basic unit of the music language, the beat unit is taken as the beat layer, and above the beat layer, there are also the measure layer, the sentence layer and the whole song layer. According to the division of the layers, the VEM-sync vector sequence is correspondingly synchronized into the beat layer, the measure layer, the sentence layer and the whole song layer. It should be noted that the physical and time sequence division method according to the measure layer, the sentence layer and the whole song layer is only one of the division levels of the present application, and is not the only one. The present application also includes the division method according to the emotion, the logic and the space.
[0143] (4) The emotion synchronization function VEM-sync. The so-called "synchronization" means that the VEM-sync vector corresponds to the corresponding VEM-Token according to the synchronization pointer. The so-called "emotional synchronization" means that the emotion description in the VEM-sync vector is synchronized with the VEM-Token. The so-called "layered fusion" of the present application means that the emotion analysis result of the whole song is obtained according to the "layered" analysis and "fusion" according to the layers.
[0144] (5) It should be noted that the VEM-sync vector is a quantity with a module length and a direction, which can be represented as a matrix or an array in mathematics. The spatial direction information has been included in the data structure. Therefore, in the present application, the VEM-sync vector, the vector and the array are regarded as the same concept and are not distinguished.
[0145] 2. VEM-Token vocal emotion multi-modal model
[0146] On the basis of the foregoing basic scheme, the present application in the VEM-Token vocal emotion multi-modal model includes but is not limited to one or a combination of the following beat capturing and alignment steps or methods:
[0147] ST110: The time measurement of the VEM-Token sequence is in seconds or beat sequence numbers, and the length starts from 0 to the end of the song.
[0148] ST120: The synchronization pointer includes a single synchronization pointer pointing to and aligning with the VEM-Token beat start time, or a double synchronization pointer pointing to the VEM-Token beat start and end times respectively.
[0149] ST130: The length of the data structure of the synchronization content includes but is not limited to fixed length and variable length.
[0150] ST140: The VEM classification refers to the classification of emotions, including but not limited to independent emotions, opposite emotion pairs, and related opposite emotion groups, and one or a combination of one-dimensional unidirectional, one-dimensional bidirectional and multi-dimensional bidirectional VEM coordinate systems are established in turn.
[0151] ST150: The VEM modalities include, but are not limited to, one or a combination of lyrics, singing, accompaniment, style, music, accompaniment instruments, video, and images.
[0152] ST160: The style includes, but is not limited to, one or a combination of folk song singing, popular song singing, Western song singing, pop song singing, original song singing, and opera singing.
[0153] ST170: The emotion name is one or a combination of, but is not limited to, joy, sadness, anger, fear, disgust, surprise, calm, anticipation, trust, love, hate, affection, and hatred in the VEM modalities, and the emotion value is set to be one or a combination of, but is not limited to, a percentage or a numerical value for the emotion metric.
[0154] ST180: The VEM library includes, but is not limited to, at least the VEM modalities, the style, the emotion name, and the emotion value, which are generated by supervised learning, reinforcement learning of typical song guidance by human vocal music experts, and deep learning of the VEM-Token vocal emotion multi-modal model.
[0155] In the present application, emotions are expressed and modeled using vectors with lengths and directions, and an emotion is represented by an emotion name and an emotion value. Emotions also include classification according to relationships, including at least independent emotions, opposite emotion pairs, and associated opposite emotion groups, and a VEM coordinate system is established in one or a combination of, but is not limited to, one-dimensional single-direction, one-dimensional bidirectional, and multi-dimensional bidirectional.
[0156] As shown in Figure 2 , this is a schematic diagram of emotion vectors, which is analyzed as follows:
[0157] (1) For the associated opposite emotion groups, such as love, hate, affection, hatred, joy, and sorrow, here, love, hate, affection, hatred, joy, and sorrow are 3 groups of opposite emotions, and there is a certain association between the emotion groups. In order to facilitate mathematical analysis, the 3 groups of emotions are included in a three-dimensional coordinate system in an orthogonal manner, Figure 2 is an example. It should be noted that according to the theory of psychology, there are other modeling methods between emotion groups, and the present application only shows one emotion group modeling method, and potential users of the present patent application can use other modeling methods.
[0158] (2) Figure 2In the middle, the emotion value of a certain E point of the three emotion groups of love, hate, emotion, enemy, joy and worry is marked by E(x, y, z) of a three-dimensional coordinate system such as a three-dimensional orthogonal coordinate system XYZ. It should be noted that since the three groups of emotions are related to each other (orthogonal relationship when added), the E point on the X, Y and Z coordinate axes generates projection values x, y and z, and also has geometric and psychological significance, and has actual follow-up calculation value for establishing a mathematical model.
[0159] It should be noted that the VEM library is based on the existence before the application of the present application, and its early stage is completed by human vocal music experts, specifically after collecting some typical songs classified according to vocal music style and VEM mode, supervised learning and reinforcement learning of artificial intelligence, and then generating results. The later stage is to perform VEM-Token vocal emotion multi-modal model by artificial intelligence, and then merge with the previous results to generate.
[0160] 3. Basic steps of synchronization content
[0161] On the basis of the foregoing scheme, the present application in the basic step of synchronization content includes but is not limited to one or more of the following steps or methods:
[0162] ST210: The beat attribute is calculated and obtained by using the VEM-Token vocal emotion multi-modal model, or is collected and obtained from the score of the song file. The beat attribute specifically includes but is not limited to: a measure is composed of more than one beat, the time unit uses the number of beats per minute, the beat type includes but is not limited to strong, secondary strong, weak, and rest, and the type includes but is not limited to 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, 3 / 8 beat, 6 / 8 beat, and 7 / 8 beat.
[0163] ST220: The emotion attribute is calculated and obtained by using the VEM-Token vocal emotion multi-modal model for the VEM-Token sequence and one or more emotion names, and the emotion value of the emotion name is normalized to between -100% and +100%. The emotion weight is set, and the emotion name, the emotion value and the emotion weight are constructed into a storage unit of an emotion weight function and stored in the VEM-sync vector sequence.
[0164] ST230: The emotion weight function is defined as a weighting function for an emotion value to modify the resulting emotion value of an emotion name in different layers in a VEM-sync vector in the VEM-sync vector sequence.
[0165] ST240: The layer is set to divide the VEM-sync vector sequence of the song file into beat layer, measure layer, sentence layer and whole song layer, and the emotion weight function is applied layer by layer to obtain the corresponding weighted emotion value.
[0166] ST250: According to the emotion change of the whole song file, the emotion synchronous function of the whole song file is calculated and obtained by hierarchical fusion according to the emotion weight function.
[0167] It should be noted that, in addition to the vertical division method from the bottom (i.e., the beat layer) to the top (i.e., the section layer, the sentence layer, and the whole song layer in turn), there is also a horizontal division method according to, for example, emotion name, instrument name, etc. The transmission of the emotion attribute includes vertical transmission and horizontal transmission, and the horizontal transmission can be further divided into forward transmission, backward transmission, and bidirectional transmission.
[0168] 4. Multi-layer weighted scanning step
[0169] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following steps or methods in the multi-layer weighted scanning step of the emotion weight function:
[0170] ST410: Obtain the emotion attribute of the beat layer of the song file in units of beats.
[0171] ST420: Add the section layer to the VEM-sync vector sequence in units of sections, take the average of the emotion values of the same emotion name in the section layer, and store each type of emotion and the corresponding emotion average value in the section layer.
[0172] ST430: Add the sentence layer to the VEM-sync vector sequence in units of sentences, accumulate the emotion values of the same emotion name in the sentence layer, extract the style weight from the VEM library according to the style, or select the style weight from the user, obtain the emotion value after the emotion name weight calculation by using the weighted algorithm, and store each type of emotion and the corresponding weighted emotion average value in the sentence layer.
[0173] ST440: Add the whole song layer to the VEM-sync vector sequence in units of whole songs, sort the emotion values of the emotion name in the sentence layer from large to small, select the first part of the emotion name or the user-specified emotion name as the high-weight emotion, wherein the first part is less than 40% according to the statistical results in the VEM library or selected by the user, store the high-weight emotion and the corresponding emotion value in the whole song layer, output as the emotion weight function, and store in the VEM library.
[0174] As Figure 3 This is a multi-layer scanning synchronization content diagram from bottom to top, which is analyzed as follows:
[0175] (1) Figure 3The 4 / 4 beat song is divided into four levels from bottom to top, namely, the beat level, the measure level, the sentence level and the whole song level. Among them, the VEM-Token vocal emotion multi-modal model is used to calculate the beat division and obtain all emotion names and emotion values of the beat, which are included in the corresponding VEM-sync vector sequence.
[0176] (2) Since the song is a 4 / 4 beat song, the emotion name and emotion value of every four beats are included in a measure layer. After the emotion name is summarized, the emotion value is calculated and stored in the measure layer.
[0177] (3) The so-called sentence layer refers to the sentence divided according to the lyrics in the song, and each sentence includes several measures. The emotion value included in the sentence layer is obtained by querying the VEM library according to the style or selected by the user.
[0178] (4) All the sentence layers are fused into the whole song emotion sequence. It should be noted that in the whole song layer, the corresponding emotion sequence is usually calculated according to the order of the sentence layer, that is, the emotion of the whole song layer is not a single value, but a function that changes with the beat time flow according to the emotion of the song. Therefore, the present application refers to it as an emotion function.
[0179] 5. Steps of recurrent neural network
[0180] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following steps or methods in the step of the emotion weight function of the recurrent neural network:
[0181] ST510: Establish a recurrent neural network RNN to realize a memory network model, realize the forward propagation of emotion, and input the VEM-sync vector emotion value of the previous beat as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value for fusion calculation, so as to enhance the emotion representation of the upper and lower beats and obtain the current emotion value.
[0182] ST520: Obtain the emotion name and emotion value of the VEM-sync vector, mark the VEM-sync vector with beat sequence number t as the hidden history state VEM-sync (t) of the current beat, mark the VEM-sync vector with beat sequence number t-1 as the hidden history state VEM-sync (t-1) of the previous beat, and t starts from 2 until the end of the song.
[0183] ST530: Set the calculation formula of the memory network model as:
[0184] VEM-SYNC(t) = f(W_x ▪ VEM-sync(t) + W_h ▪ VEM-sync(t-1) + b),
[0185] where VEM-SYNC(t) is the hidden state, f is the activation function, W_x, W_h are the trainable emotion weight matrices, b is the bias vector, the output of VEM-SYNC(t) is stored in the VEM library as the emotion weight function.
[0186] This is a typical emotion forward propagation scheme based on recurrent neural network RNN. By modifying the recurrent network, a vertical propagation scheme based on convolutional neural network CNN can be supported.
[0187] 6. Long short-term memory network steps
[0188] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following steps or methods in the long short-term memory network steps of the emotion weight function:
[0189] ST610: Establish a long short-term memory network LSTM to realize a long-term memory network model that exceeds the previous beat, realize the forward propagation of emotion, and input the VEM-sync vector emotion value of the forgotten beat or forgotten emotion as the hidden history state of the current beat when calculating the current beat VEM-sync vector, combine the current emotion value to calculate the fusion, so as to enhance the emotion representation of the upper and lower beats and obtain the current emotion value.
[0190] ST620: The forgotten beat is an unimportant beat, including but not limited to more than one local accompaniment, pass door, singing, breath in a song file, or a beat selected by a user, and the forgotten emotion also includes an unimportant emotion name and its emotion value, or an emotion name and its emotion value selected by a user.
[0191] ST630: The output of the memory network model is the emotion weight function, which is stored in the VEM library.
[0192] It should be noted that in the VEM library, there are contents previously annotated by human vocal music experts, including the labels of the memory network model, especially for some special songs, the audience emotion sound of the singing and performance scene environment is also recorded and labeled.
[0193] 7. Self-attention mechanism steps
[0194] On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following model steps or methods in the self-attention mechanism steps of the emotion weight function:
[0195] ST710: Obtain the emotion name and emotion value of the VEM-sync vector of the song file, and mark the VEM-sync vector with the beat sequence number t as VEM-sync(t).
[0196] ST720: Establish a Transformer encoder to realize the forward propagation, backward propagation and omnidirectional propagation of emotions in the VEM-sync vector sequence of all beats of the song file through the self-attention mechanism, calculate the global dependency between all beat pairs in the VEM-sync vector sequence, and output a set of emotion value sequences enhanced by global up and down VEM-Token sequence weights.
[0197] ST730: The Transformer encoder includes but is not limited to one or more encoder layers, and each encoder layer further includes but is not limited to a multi-head self-attention layer and a feedforward neural network layer. Through residual connection, layer normalization, input embedding and position encoding, the update and output of the emotion weight function for all VEM-sync vector sequences are completed, and the position encoding is obtained from the beat sequence number.
[0198] ST740: Input all VEM-sync vector sequences of the song file into the Transformer encoder at one time, and set one or more emotion names and emotion values that need to be focused on to access the input end of the multi-head self-attention layer, where the input end is v. In the VEM-sync vector, add a multi-head self-attention layer and a feedforward neural network layer, and the Transformer encoder stores the calculation results in the multi-head self-attention layer.
[0199] ST750: Set a dynamic mechanism selection to access the emotion name and emotion value of the multi-head self-attention layer.
[0200] ST760: Map the emotion name and emotion value of each beat to the three spaces including but not limited to query Query, key Key, and value Value through a learnable weight matrix. On the beat sequence, calculate the dot product of Query and Key to obtain the association weight between beats, and use this association weight to weight sum all Value, thereby obtaining the emotion weight information of the song file.
[0201] ST770: For the emotion name and emotion value not included in the input of the multi-head self-attention layer, select the top N emotion names, emotion values and beat numbers with the largest emotion values through the RAG step of setting the prompt word and enhancing the generation by retrieval, and incorporate them into the emotion weight function to prevent omission, where the value of N is determined according to the song file or determined by the user.
[0202] ST780: output the result stored in the multi-head self-attention layer as an emotion weight function, and store it in the VEM library.
[0203] As shown in Figure 4 , this is a schematic diagram of the self-attention mechanism step of the emotion weight function. The analysis is as follows:
[0204] (1) Figure 4 It includes three parts, namely the emotion synchronization function VEM-sync, the VEM-Token sequence, and the video background. Among them, the VEM-Token sequence is obtained by using the VEM-Token vocal emotion multi-modal model, dividing and aligning the beats, and it needs to be noted that the VEM-Token sequence here is actually a spectrum format file, only with beat information added. The video background is the MV (Music Video, music video, including moving pictures and still pictures) in the song file.
[0205] (2) The emotion synchronization function VEM-sync is a fixed-length or variable-length array, also known as a matrix. Among them, the synchronization pointers point to the VEM-Token sequence and the beat part of the video background divided according to the beat respectively; the VEM-sync vector includes a two-dimensional array composed of multiple emotion names, emotion values, and emotion weights; the emotion names and emotion values are calculated by using the VEM-Token vocal emotion multi-modal model on the corresponding VEM-Token sequence.
[0206] (3) The emotion synchronization function VEM-sync here includes being composed according to the Transformer encoder, using the multi-head self-attention mechanism, and also includes being composed according to the RAG, using the enhanced retrieval generation mechanism.
[0207] (4) It needs to be noted that the division of labor or adoption or not of the multi-head self-attention layer and the RAG layer can be determined by the user according to the attributes of the emotion.
[0208] Figure 5 is a schematic diagram of emotion weight synchronization. It not only has reference value for this step, but also has reference value for the aforementioned multi-layer weighted scanning step. In Figure 5 , it can be divided into the left emotion attribute initial matrix including emotion name, emotion value and emotion weight, the middle emotion synchronization function, and the right emotion attribute result matrix including emotion name, emotion value and emotion weight. Here, the emotion synchronization function will calculate the emotion weight in the initial matrix through this step, and store the result in the result matrix.
[0209] 8. RAG step of enhanced generation
[0210] On the basis of the foregoing scheme, the step of outputting the emotion weight function further includes, but is not limited to, one or a combination of the following steps or methods:
[0211] ST810: Obtain the tempo, emotion name and emotion value of the VEM-sync vector sequence of the song file, and mark the VEM-sync vector with the tempo t as the VEM-sync(t) array, wherein the value range of t includes more than one paragraph of the song file.
[0212] ST820: According to the VEM-sync(t) array, retrieve the VEM library, and obtain more than one VEM-sync vector sequence of K-RAG emotion values similar to the local emotion name or all emotion names in the VEM-sync(t) array as a reference VEM-sync(t) array sequence, and the similarity is determined by the user.
[0213] ST830: Take the VEM-sync(t) array and the reference VEM-sync(t) array as the input end of the multi-head self-attention mechanism of the Transformer encoder, adopt the self-attention mechanism, calculate the cross-attention weight of the VEM-sync(t) array on the reference VEM-sync(t) array, and then perform weighted fusion to obtain the fusion result VEM-sync(t) array.
[0214] ST840: The fusion result VEM-sync(t) array is used to verify or correct the preliminary emotion analysis result calculated by the multi-head self-attention network.
[0215] Referring to Figure 5 and the description.
[0216] It should be noted that in the step ST830, "calculating the cross-attention weight of the VEM-sync(t) array on the reference VEM-sync(t) array and then performing weighted fusion", which includes vertically merging or horizontally merging the two groups of matrices.
[0217] 9. The step of outputting the emotion weight function
[0218] On the basis of the foregoing scheme, the step of outputting the emotion weight function further includes, but is not limited to, one or a combination of the following steps or methods:
[0219] ST910: According to the emotion weight function of the entire VEM-sync vector sequence of the song file, convert the emotion score output into a written format, a digital format or a MIDI communication interface format.
[0220] ST920: Establish the format of communication with the Agent side, output the emotional score to the Agent side according to the mood weight function of the entire VEM-sync vector sequence of the song file.
[0221] ST930: Establish an NLP-token-oriented interface according to the criteria of NLP-token words in the natural language model NLP-token, including prompt word interface, context protocol MCP interface and AI interface.
[0222] ST950: For song files with video and background pictures synchronized in time, establish a bidirectional picture pointer one pointing to and aligning the beat starting point and picture synchronization point according to the tempo and beat starting point of the VEM-Token sequence of the song file.
[0223] ST960: Add picture pointer two in the VEM-sync vector sequence, which points to and aligns the picture synchronization point according to the tempo and beat starting point of the VEM-Token sequence.
[0224] ST970: Establish the synchronization and alignment relationship between the VEM-sync vector sequence and the video and background pictures through the tempo and beat starting point of the VEM-Token sequence, the synchronization pointer, the picture pointer one and the picture pointer two.
[0225] ST980: According to user demand, output fusion information from the VEM-sync vector sequence to the video and background pictures, including but not limited to lyrics, song score and information or actions in the VEM-sync vector sequence required by the user.
[0226] ST990: According to user demand, output fusion information from the VEM-sync vector sequence to include but not limited to 3D animation expression action, language action and body action, to drive include but not limited to 3D animation and expression action, language action and body action, to make expressions and body actions corresponding to the emotional vector in the VEM-sync vector sequence.
[0227] ST9A0: Support Model Context Protocol MCP (Model Context Protocol), including MCP Host, MCPClient, MCP Server interface, provide Function Calling function. In order to use the invention with other large models or AI systems.
[0228] Figure 6 is a full song mood synchronization function fusion diagram. The analysis is as follows:
[0229] (1) Figure 6In the middle, 601 is the emotional projection slice of the nodes in the full curve layer, 602 is the emotional synchronization function vector, and 603 is the beat line.
[0230] (2) The emotional projection slice is a diagram of the emotional synchronization function at an instant. Since the emotional synchronization function is the instantaneous projection of the multi-dimensional emotional vector on the beat time axis, for the convenience of explanation, the emotional projection slice is drawn as a slice containing only love, hate, emotion, and hatred. In fact, it should be the instantaneous projection of the synthesis or fusion of all emotional categories.
[0231] (3) For the full song, the emotional synchronization function vector is a kind of high-dimensional continuous variable according to the full song time.
[0232] 10. The step of intellectual property record
[0233] On the basis of the foregoing scheme, the present application also includes but is not limited to the steps of intellectual property record and management, specifically including one or more combinations of the following steps or methods:
[0234] STA10: Insert the mark vector space CV or interface of the intellectual property record in the whole VEM-sync vector sequence of the song file, used to store the copyright vector.
[0235] STA20: The copyright vector includes but is not limited to copyright holder information, version number, authorization information, encryption and decryption mode.
[0236] STA30: Including encryption and decryption engine, including but not limited to elliptic algorithm, public key, private key mode.
[0237] STA40: Managed by block chain and cloud mode.
[0238] It should be noted that one of the features that distinguishes the present application from other methods is the VEM-sync vector sequence with emotional synchronization function, which provides a storage space outside the song file and its VEM-Token sequence, therefore, using the storage space of the VEM-sync vector sequence to process the storage of intellectual property does not affect the song file itself, and is also conducive to the storage of encryption and decryption.
Claims
1. A method for hierarchical fusion of VEM-Token emotion synchronization functions, characterized in that, Comprising: ST100: using a VEM-Token vocal emotion multi-modal model to divide a song file into a VEM-Token sequence in units of beats, and setting an emotion synchronization function as VEM-sync, wherein the VEM-sync includes a VEM-sync vector sequence corresponding to and aligned with the VEM-Token sequence, and the VEM-sync vector includes synchronization content and a synchronization pointer pointing to a corresponding VEM-Token beat; ST200: according to the VEM-Token sequence, calculating the synchronization content of the VEM-sync vector sequence, and the synchronization content including beat attributes and emotion attributes, wherein the emotion attributes include one or more combinations of emotion names, emotion values and emotion weights; ST300: the emotion weight function includes: using hierarchical processing of a song file to obtain a corresponding VEM-sync vector sequence, using one or a combination of multi-layer weighted scanning, recurrent neural network, long short-term memory network, self-attention mechanism and retrieval enhanced generation for forward propagation, backward propagation and omnidirectional propagation to obtain an emotion weight function, and applying the emotion weight function layer by layer, and layer-by-layer fusion to obtain an emotion synchronization function of the song file.
2. The method according to claim 1, characterized in that, ST100 includes: ST110: the time measurement of the VEM-Token sequence is seconds or beats, and the length starts from 0 to the end of the song; ST120: the synchronization pointer includes: a single synchronization pointer pointing to and aligned with the VEM-Token beat start time, or a double synchronization pointer pointing to the VEM-Token beat start and end times respectively; ST130: the length of the data structure of the synchronization content includes fixed length and variable length; ST140: the VEM classification refers to the classification of emotions, including independent emotions, opposite emotion pairs, and related opposite emotion groups, and a VEM coordinate system including one or a combination of one-dimensional unidirectional, one-dimensional bidirectional and multi-dimensional bidirectional is established in turn; ST150: the VEM modalities include one or a combination of lyrics, singing, accompaniment, style, music, accompaniment instruments, video and image; ST160: the style includes one or a combination of national song singing method, popular song singing method, western song singing method, pop song singing method, original song singing method, and opera singing method; ST170: the emotion name is one or a combination of joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, emotion and hatred in the VEM modalities, and the emotion value is set as a percentage or a numerical value for emotion measurement; ST180: the VEM library at least includes VEM modalities, styles, emotion names and emotion values, which are generated by VEM-Token vocal emotion multi-modal model through supervised learning, reinforcement learning by human vocal expert guidance, and deep learning.
3. The method according to claim 2, characterized in that, The basic steps of the synchronization content include: ST210: The beat attribute is calculated by the VEM-Token vocal music emotion multi-modal model or obtained from the music file score collection. The beat attribute specifically includes: time signature, measure, beat type, and type, wherein the measure is composed of more than one beat, the timing unit uses the number of beats per minute, the beat type includes strong, secondary strong, weak, and rest, and the type includes 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, 3 / 8 beat, 6 / 8 beat, and 7 / 8 beat; ST220: The emotion attribute is calculated by the VEM-Token vocal music emotion multi-modal model for the VEM-Token sequence and more than one emotion name. The emotion value of the emotion name is calculated and normalized to between -100% and +100%. The emotion weight is set, and the time signature, emotion name, emotion value, and emotion weight are combined to form an emotion weight function and stored in the VEM-sync vector sequence; ST230: The emotion weight function is defined as a weighting function for an emotion value to modify the resulting emotion value of an emotion name in different layers in a VEM-sync vector in the VEM-sync vector sequence; ST240: The layer is set as the VEM-sync vector sequence of the music file, which is divided into beat layer, measure layer, sentence layer, and whole song layer. The emotion weight function is applied layer by layer to obtain the corresponding weighted emotion value; ST250: According to the emotion change of the whole music file, the emotion weight function is calculated by the hierarchical fusion method to obtain the emotion synchronization function of the whole music file.
4. The method according to claim 3, characterized in that, The emotion weight function includes a multi-layer weighting scanning step, specifically including: ST410: The beat unit is used to obtain the beat layer emotion attribute of the music file; ST420: The measure unit is used to add the measure layer to the VEM-sync vector sequence. The average value of the emotion values of the same emotion name in the measure layer is taken, and the emotion of each type and the corresponding emotion average value are stored in the measure layer; ST430: The sentence unit is used to add the sentence layer to the VEM-sync vector sequence. The emotion values of the same emotion name in the sentence layer are accumulated. The style weight is extracted from the VEM library according to the style, or the style weight is selected by the user. The weighted emotion value of the emotion name is obtained by using the weighting algorithm. The emotion of each type and the corresponding weighted emotion average value are stored in the sentence layer; ST440: The whole song unit is used to add the whole song layer to the VEM-sync vector sequence. The emotion values of the emotion names in the sentence layer are sorted from large to small. The first part of the emotion name or the user-specified emotion name is selected as the high-weight emotion. The first part is determined according to the statistical result in the VEM library or selected by the user. The high-weight emotion and the corresponding emotion value are stored in the whole song layer. The output is used as the emotion weight function and stored in the VEM library.
5. The method according to claim 3, characterized in that, The emotion weight function includes the steps of a recurrent neural network, specifically including: ST510: Establish a recurrent neural network (RNN) to implement a memory network model, implement forward propagation of emotions, and input the emotion value of the VEM-sync vector of the previous beat as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value to calculate the fusion, to enhance the emotion representation of the upper and lower beats, and obtain the current emotion value; ST520: Obtain the emotion name and emotion value of the VEM-sync vector, mark the VEM-sync vector with beat number t as the hidden history state VEM-sync(t) of the current beat, and mark the VEM-sync vector with beat number t-1 as the hidden history state VEM-sync(t-1) of the previous beat, t starts from 2 until the end of the song; ST530: Set the calculation formula of the memory network model as: VEM-SYNC(t) = f(W_x·VEM-sync(t) + W_h·VEM-sync(t-1) + b), where VEM-SYNC(t) is the hidden state, f is the activation function, W_x and W_h are trainable emotion weight matrices, and b is the bias vector. The VEM-SYNC(t) output is stored as an emotion weight function in the VEM library.
6. The method according to claim 3, characterized by The emotion weight function includes the steps of a long short-term memory network, specifically including: ST610: Establish a long short-term memory network (LSTM) to implement a longer-term memory network model than the previous beat, implement forward propagation of emotions, and input the emotion value of the VEM-sync vector of the forgotten beat or forgotten emotion as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value to calculate the fusion, to enhance the emotion representation of the upper and lower beats, and obtain the current emotion value; ST620: The forgotten beat is an unimportant beat, including one or more local accompaniment, pass door, singing, breath in the song file, or the beat selected by the user, and the forgotten emotion is an unimportant emotion name and its emotion value, or the emotion name and its emotion value selected by the user; ST630: The output of this memory network model is the emotion weight function, which is stored in the VEM library.
7. The method according to claim 3, characterized in that, The emotion weight function includes the steps of a self-attention mechanism, specifically including: ST710: Obtain the beat number, emotion name, and emotion value of the VEM-sync vector of the song file, and mark the VEM-sync vector with beat number t as VEM-sync(t); ST720: Establish a Transformer encoder to implement forward propagation, backward propagation, and omnidirectional propagation of emotions in the VEM-sync vector sequence of all beats of the song file through a self-attention mechanism, calculate the global dependency relationship between all beat pairs in the VEM-sync vector sequence, and output a set of emotion value sequences enhanced by global upper and lower VEM-Token sequence weights; ST730: The Transformer encoder includes one or more encoder layers, each of which further includes a multi-head self-attention layer and a feed-forward neural network layer, which updates and outputs the emotion weight function for the entire VEM-sync vector sequence through a residual connection, layer normalization, input embedding, and position encoding, where the position encoding is taken from the beat number; ST740: The Transformer encoder is inputted with the entire VEM-sync vector sequence of the song file, and one or more emotion names and emotion values that need to be focused on are set to access the input end of the multi-head self-attention layer, and the VEM-sync vector is added to include the multi-head self-attention layer and the feed-forward neural network layer, and the Transformer encoder stores the calculation result in the multi-head self-attention layer; ST750: Set a dynamic mechanism selection to access the emotion name and emotion value of the multi-head self-attention layer; ST760: Map each beat's emotion name and emotion value to the Query, Key, and Value spaces through a learnable weight matrix, calculate the dot product of Query and Key on the beat sequence to obtain the association weight between beats, and use this association weight to weight sum all values Value to obtain the song file emotion weight information; ST770: For emotion names and emotion values not included in the multi-head self-attention layer input, select the top N emotion names, emotion values, and beat numbers with the largest emotion values by setting a prompt word and through the RAG step of retrieval augmented generation, and incorporate them into the emotion weight function to prevent omissions, where the value of N is determined according to the song file or by the user; ST780: Output the result stored in the multi-head self-attention layer as the emotion weight function and store it in the VEM library.
8. The method according to claim 7, characterized in that, The emotion weight function includes the RAG step of retrieval augmented generation, which specifically includes: ST810: Obtain the beat number, emotion name, and emotion value of the VEM-sync vector sequence of the song file, and mark the VEM-sync vector with beat number t as the VEM-sync(t) array, where the value of t includes one or more paragraphs of the song file; ST820: Retrieve the VEM library according to the VEM-sync(t) array to obtain one or more VEM-sync vector sequences with the top K-RAG emotion values similar to the local or all emotion names in the VEM-sync(t) array as reference VEM-sync(t) array sequences, where the similarity is determined by the user; ST830: Take the VEM-sync(t) array and the reference VEM-sync(t) array as the input end of the multi-head self-attention mechanism of the Transformer encoder, and use the self-attention mechanism to calculate the cross-attention weight between the VEM-sync(t) array and the reference VEM-sync(t) array, then perform weighted fusion to obtain the fusion result VEM-sync(t) array; ST840: The fusion result VEM-sync(t) array is used to verify or correct the preliminary emotion analysis result calculated by the multi-head self-attention network.
9. The method of any of claims 4, 5, 6, 7, 8, wherein, The emotion weight function includes the following steps: ST910: According to the emotion weight function of the entire VEM-sync vector sequence of the song file, convert the output into a written format, a digital format, or a MIDI communication interface format emotion score output; or, ST920: Establish a format for communication with the Agent end, and output the emotion score to the Agent end according to the emotion weight function of the entire VEM-sync vector sequence of the song file; Or, ST930: According to the standard of NLP-token word units of the natural language model, establish an NLP-token interface, including a prompt word interface, a context protocol MCP interface, and an AI interface; or, ST950: For a song file with a video and a background picture that are synchronized in time, according to the tempo and beat start point of the VEM-Token sequence of the song file, establish a bidirectional picture pointer one that points to and aligns the beat start point and the picture synchronization point, respectively; ST960: Add a picture pointer two in the VEM-sync vector sequence, which points to and aligns the picture synchronization point according to the tempo and beat start point of the VEM-Token sequence; ST970: Establish a synchronization and alignment relationship between the VEM-sync vector sequence and the video and background picture through the tempo and beat start point of the VEM-Token sequence, the synchronization pointer, the picture pointer one, and the picture pointer two; ST980: According to user requirements, output fusion information from the VEM-sync vector sequence to the video and background picture, including lyrics, song score, and information or actions in the VEM-sync vector sequence that meet user requirements; ST990: According to user requirements, output fusion information from the VEM-sync vector sequence to include 3D animation expression actions, language actions, and body actions to drive 3D animation and expression actions, language actions, and body actions to make expressions and body actions corresponding to the emotion vectors in the VEM-sync vector sequence; ST9A0: Support the model context protocol MCP, including the MCP Host, MCP Client, and MCP Server interfaces, and provide function calling functions.
10. The method according to claim 9, characterized in that, The steps include intellectual property records: STA10: Insert the copyright vector space CV or interface of the intellectual property record in the entire VEM-sync vector sequence of the song file to store the copyright vector; STA20: The copyright vector includes copyright holder information, version number, authorization information, and encryption and decryption methods; STA30: Include encryption and decryption engines, including elliptic algorithms, public key, and private key modes; STA40: Managed using blockchain and cloud mode.
Citation Information
Patent Citations
VEM-Token beat capture and alignment model construction method
CN120748450A
Speech and emotion synchronous recognition method based on neural network
CN108806667A
Digital modulation and demodulation system and method based on hard synchronization
CN117938611A