VEM-Token emotion synchronization function hierarchical fusion method

By employing a hierarchical fusion method based on the VEM-Token emotion synchronization function, the shortcomings of the NLP-Token model in parsing non-textual information modalities are addressed, enabling high-dimensional continuous quantitative description and dynamic recognition of song emotions, and supporting intelligent agents for music AI applications.

CN120913602AActive Publication Date: 2025-11-07GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD

Patent Information

Application Number
CN202511428558.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-01
Publication Date
2025-11-07
Estimated Expiration
2045-10-01

AI Technical Summary

Technical Problem

Existing technologies that rely on NLP-Token models cannot accurately and in real time parse non-textual information modalities such as singing, speech, and musical emotions. This results in inaccurate emotion classification boundaries, difficulty in describing spatial correlations, and difficulty in capturing the continuity of emotion changes over time, creating an illusion of large model outputs.

Method used

We employ a VEM-Token emotion synchronization function hierarchical fusion method. By using the VEM-Token vocal emotion multimodal model to segment song files using beats as units, and combining multi-layer weighted scanning, recurrent neural networks, long short-term memory networks, and self-attention mechanisms, we calculate the emotion synchronization function of the song files to achieve high-dimensional description and dynamic recognition of emotions.

Benefits of technology

It achieves high-dimensional continuous quantitative description of song emotions, avoids the bias of text description, supports real-time emotion analysis, drives applications such as sound-to-text/emotion-to-text and sound-to-animation/emotion-to-animation, and broadens the intelligent agent of AI music applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913602A_ABST
    Figure CN120913602A_ABST
Patent Text Reader

Abstract

A VEM-Token emotion synchronization function hierarchical fusion method is different from a traditional NLP-Token method, an emotion synchronization function VEM-sync is innovated for the first time, the emotion synchronization function VEM-sync is synchronized with a VEM-Token sequence of music beats, multiple high-dimensional emotions are directly described by adopting mathematical languages, and therefore deviation of discretized natural language characters on description of a high-dimensional emotion analog quantity function is avoided, and the emotion synchronization effect is improved. The method comprises the steps of defining VEM-sync and synchronous content, defining rhythm attributes, emotion attributes and emotion functions, adopting one or combination of multi-layer weighted scanning, a recurrent neural network, a long and short-term memory network, a self-attention mechanism and an RAG network generated by retrieval enhancement, and performing hierarchical fusion calculation to output a time emotion function of a vocal music file. According to the phonetic function or the emotional function, the effects of emotional texts, emotional expressions, emotional languages and emotional multi-dimensional animations are directly driven by crossing discrete text tokens, a model context protocol (MCP) and a function calling function are supported, and copyright management and encryption and decryption or interfaces are provided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular to the model construction and processing of AI music, singing, voice, recitation, disease patient voice change and AI intelligent agent Agent, and especially to an innovative emotion synchronization function design and operation method. Unlike the traditional NLP-token word element system under the dependence on Internet text for music, singing and voice emotion analysis, the present application first proposes a new system based on VEM-Token vocal music voice emotion recognition, and adopts a VEM-Token emotion synchronization function hierarchical fusion method to reconstruct the analysis model system. BACKGROUND

[0002] In the field of artificial intelligence, so far the decomposition of information is still based on the natural language word element division method of NLP-Token (Natural-Language-Processing Token, natural language processing word element). If it is based on text information modal analysis and application, NLP-Token has natural advantages, and large models have learned all the books and web pages based on human text, and have memorized them in the large model. However, if it is for non-text information modal, such as singing, voice, recitation emotion, music style and other forms of information modal, today's large model still looks for the text description learned from books and web pages in the past in the model memory. Through these text descriptions, "interpretation" and "understanding" are obtained, that is, the information modal based on NLP-Token word element is still adopted. Since we cannot know where the large model obtains those corpus, and cannot predict the correctness of the corpus, the "illusion" of the large model cannot be avoided. In addition, many information modal in nature cannot be accurately described directly by language and text. For example, human emotions and even animal emotions, according to the classification of psychology, there are at least joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, emotion, and hatred. By carefully examining these emotions, we will find that: (1) The classification boundary of emotion cannot be accurately described by text. For example, it is difficult to accurately define the classification boundary between "joy" and "love", and how to define this ambiguity and convertibility in text description? (2) There is a spatial correlation between emotions, which is difficult to describe in text. Even if we agree with the above psychological classification, for example, "fear", "disgust" and "surprise" are different categories, but how do they affect and relate to each other? Although NLP-Token can measure the angle and modulus of the vector in the high-dimensional coordinate system based on the dimension of the emotion category, what if the coordinate system itself has a problem? (3) Emotion is a continuous analog function that changes over time, and the text description can only be a slice of this function, and cannot be described in real time. For example, when a song describes "sadness" and "hate", the singer's emotions will fluctuate as the rhythm flows, and different singers, different environments, and even different times will have different fluctuations. For example, Xianglin'ao repeatedly tells the story of Mao being eaten by a wolf, sometimes with sad tears, and sometimes with calm tears.

[0003] By analogy, we can easily find that even if the knowledge learned by the large model is more, it is also impossible to track the emotions in real time and output accurate "explanation" and "understanding". The fundamental reason for this situation is the NLP-Token model itself.

[0004] The inventors' team first proposed a new model design, which includes an already granted Chinese invention patent "VEM-Token vocal emotion multimodal tokenization song and accompaniment deep learning method, CN120126506" (hereinafter referred to as "VEM-Token vocal emotion multimodal model") and an invention patent "VEM-Token beat capture and alignment model construction method, 202511249168.0" (hereinafter referred to as "VEM-Token beat model") that is under review. The music beat is successfully segmented as an information token, i.e. VEM-Token (Vocal-Emotion-Multimodal Token), which is different from the traditional NLP-token method, and a new emotion synchronization function VEM-sync is innovated, which is synchronized with the VEM-Token sequence of the music beat, and directly describes high-dimensional emotions in mathematical language, thereby avoiding the deviation and illusion of natural language text in describing high-dimensional emotion functions. The invention includes VEM-sync and synchronous content definition, beat attribute, emotion attribute, emotion function definition, one or combination of multi-layer weighted scanning, recurrent neural network, long short-term memory network, self-attention mechanism, retrieval enhancement generation RAG network, hierarchical fusion calculation output vocal file time emotion function, according to this sound function or emotion function, directly drive sound text / emotion text, sound expression / emotion expression, sound multi-dimensional animation / emotion multi-dimensional animation, support model context protocol MCP and function calling Function Calling, provide support for copyright management and encryption / decryption or interface.

[0005] The present application is a patent pool patent of CN120126506, 202511249168.0, which can access an AI system or an independently developed application system, and develop a vocal music intelligent agent that can listen to songs and recognize music scores. Through supervised learning, the model obtains a VEM library. Due to the existence of this professional and accurate learning VEM library supervised by human experts, the possibility of "hallucination" produced by the large model in application almost disappears.

[0006] In addition, as a VEM-Token-based patent pool patent, the present application has expansion potential in supporting embodied robot facial expressions, supporting virtual digital humans, and supporting digital user self.

[0007] The VEM-Token concept mentioned in the present application is based on the basic concepts and steps defined in CN120126506, 202511249168.0 invention patent, unless otherwise emphasized or specifically defined in the present application. Among them, "vocal music files", "music files", "songs", and "music" mentioned in the present application file have the same meaning unless otherwise stated.

[0008] The present application aims to be applied to the research and development and application of current large models, such as OpenAI, DeepSeek, Google Gemini, Kimi, bean bag large model, and Wenxin Yiyang, etc. After connecting the front end and back end of some applications, various intelligent agents are formed. The present application aims to access these large models and achieve bidirectional communication, form music / vocal music-based artificial intelligence applications, and even music AI intelligent agents, providing powerful innovation and support for broadening the application of AI.

[0009] Disadvantages of prior art methods (1) The existing NLP-token-based technology analyzes transitions and describes emotions through text, but text is discrete and has clear boundaries, while emotions are not only continuous analog quantities and fuzzy concepts, but also time functions that change over time. Therefore, using NLP-token to describe emotions is not a good approach, which will lose a lot of information details in emotions.

[0010] (2) Currently, there is no mathematical model of a vocal music / musical emotion generating function, i.e., a "sound generating function / emotion generating function" in common terms, so there is no method of generating "sound generating text / emotion generating text", "sound generating animation / emotion generating animation", and "sound generating action / emotion generating action" driven by vocal music / musical emotions. SUMMARY

[0011] The present application is a new VEM-Token emotion synchronization function hierarchical fusion method, which is based on the VEM-Token vocal emotion multi-modal model, rather than the NLP-token model, to automatically identify, manage and process vocal emotion, achieve the purpose and intention of the present application, and provide an effective preliminary basis for realizing "sound generates text / emotion generates text", "sound generates animation / emotion generates animation" and "sound generates action / emotion generates action".

[0012] The purpose and intention of the present application are achieved by using the following technical solutions and working steps: 1. VEM-Token emotion synchronization function hierarchical fusion method The VEM-Token emotion synchronization function hierarchical fusion method of the present application includes but is not limited to the following steps: ST100: Using the VEM-Token vocal emotion multi-modal model, the song file is segmented into VEM-Token sequences in units of beats, and the emotion synchronization function is set as VEM-sync, wherein VEM-sync includes but is not limited to a VEM-sync vector sequence corresponding to and aligned with the VEM-Token sequence, and the VEM-sync vector includes but is not limited to synchronization content and a synchronization pointer pointing to the corresponding VEM-Token beat.

[0013] ST200: According to the VEM-Token sequence, the synchronization content of the VEM-sync vector sequence is calculated and obtained, and the synchronization content includes but is not limited to beat attributes and emotion attributes, wherein the emotion attributes include but are not limited to one or more emotion names, emotion values and emotion weights.

[0014] ST300: The emotion weight function includes but is not limited to: using hierarchical processing of the song file to obtain the corresponding VEM-sync vector sequence, and the emotion value thereof includes but is not limited to hierarchical emotion weights, using one or a combination of multi-layer weighted scanning, recurrent neural network, long short-term memory network, self-attention mechanism and retrieval enhancement generation for forward propagation, backward propagation and omnidirectional propagation, and fusion to obtain the emotion synchronization function of the song file.

[0015] 2. VEM-Token vocal emotion multi-modal model On the basis of the foregoing scheme, the present application in the VEM-Token vocal emotion multi-modal model includes but is not limited to one or more combinations of the following steps or methods of beat capture and alignment: ST110: The time measurement of the VEM-Token sequence is in seconds or beats, and the length starts from 0 to the end of the song.

[0016] ST120: The synchronization pointer includes a single synchronization pointer pointing to the VEM-Token beat start time, or a double synchronization pointer pointing to the VEM-Token beat start time and end time respectively.

[0017] ST130: The length of the data structure of the synchronization content includes but is not limited to fixed length and variable length.

[0018] ST140: The VEM classification refers to the classification of emotions, including but not limited to independent emotions, opposite emotion pairs, and related opposite emotion groups, and in turn establishes one or a combination of VEM coordinate systems including but not limited to one-dimensional unidirectional, one-dimensional bidirectional, and multi-dimensional bidirectional.

[0019] ST150: The VEM mode includes but is not limited to one or a combination of lyrics, singing, accompaniment, style, music, accompaniment instruments, video, and image.

[0020] ST160: The style includes but is not limited to one or a combination of national song singing, popular song singing, western song singing, pop song singing, original song singing, and opera singing.

[0021] ST170: The emotion name is one or a combination of joy, sadness, anger, fear, disgust, surprise, calmness, expectation, trust, love, hate, emotion, and hatred included in the VEM mode, and the emotion value is set to one or a combination of percentage or numerical value for emotion measurement.

[0022] ST180: The VEM library includes but is not limited to at least VEM mode, style, emotion name, and emotion value, which are generated by VEM-sync vector sequence generated by supervised learning, reinforcement learning of typical song guidance by human vocal music experts, and deep learning of VEM-Token vocal emotion multi-modal model.

[0023] 3. Basic steps of synchronization content On the basis of the foregoing scheme, the present application includes but is not limited to one or more of the following steps or methods in the basic steps of synchronization content: ST210: The beat attribute is calculated and obtained by the VEM-Token vocal emotion multi-modal model, or is obtained by collecting the score of the song file, and the beat attribute specifically includes but is not limited to tempo, measure, beat type, and type, wherein the measure is composed of more than one beat, the time unit uses the number of beats per minute, the beat type includes but is not limited to strong, secondary strong, weak, and rest, and the type includes but is not limited to 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, 3 / 8 beat, 6 / 8 beat, and 7 / 8 beat.

[0024] ST220: Emotional attribute, using VEM-Token vocal emotion multi-modal model, for VEM-Token sequence and one or more emotion names, calculate the emotion value of the obtained emotion name, and normalize the emotion value to between -100% and +100%, set the emotion weight, and construct the storage unit of the emotion weight function of the beat number, emotion name, emotion value and emotion weight, and store it in the VEM-sync vector sequence.

[0025] ST230: The emotion weight function is defined as a weighting function for an emotion value to modify the resulting emotion value of an emotion name in different layers in a VEM-sync vector in the VEM-sync vector sequence.

[0026] ST240: Set the layer to the VEM-sync vector sequence of the song file, divide it into beat layer, measure layer, sentence layer and whole song layer, and apply the emotion weight function layer by layer to obtain the corresponding weighted emotion value.

[0027] ST250: According to the emotion change of the whole song file, the emotion weight function is used to calculate the emotion synchronization function of the whole song file according to the hierarchical fusion.

[0028] 4. Multi-layer weighted scanning step of emotion weight function On the basis of the foregoing scheme, the present application in the multi-layer weighted scanning step of the emotion weight function includes but is not limited to one or more combinations of the following steps or methods: ST410: Get the beat layer emotion attribute of the song file in units of beats.

[0029] ST420: Add measure layer in VEM-sync vector sequence in units of measures, take the average value of the emotion value of the same emotion name in the measure layer, and store each kind of emotion and the corresponding emotion average value in the measure layer.

[0030] ST430: Add sentence layer in VEM-sync vector sequence in units of sentences, do accumulation on the emotion value of the same emotion name in the sentence layer, extract style weight in VEM library according to style, or select style weight by user, obtain emotion value after weighted calculation of emotion name by using weighted algorithm, store each kind of emotion and the corresponding weighted emotion average value in the sentence layer.

[0031] ST440: In the VEM-sync vector sequence, add the full song layer in the full song unit, sort the emotion values of the emotion names in the sentence layer from large to small, select the emotion names of the first part or the emotion names specified by the user as high-weight emotions, wherein the first part is selected according to the statistical results in the VEM library or by the user or less than 40%, store the high-weight emotions and their corresponding emotion values in the full song layer, output as an emotion weight function, and store in the VEM library.

[0032] 5. The step of the recurrent neural network of the emotion weight function On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following steps or methods in the step of the recurrent neural network of the emotion weight function: ST510: Establish a recurrent neural network RNN to realize a memory network model, realize the forward propagation of emotions, input the emotion value of the VEM-sync vector of the last beat as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value to fuse and calculate, so as to enhance the emotion representation of the upper and lower beats, and obtain the current emotion value.

[0033] ST520: Obtain the emotion name and emotion value of the VEM-sync vector, mark the VEM-sync vector with beat number t as the hidden history state VEM-sync (t) of the current beat, mark the VEM-sync vector with beat number t-1 as the hidden history state VEM-sync (t-1) of the previous beat, and t starts from 2 until the end of the song.

[0034] ST530: Set the calculation formula of the memory network model as: VEM-SYNC (t) = f (W_x ▪ VEM-sync (t) + W_h ▪ VEM-sync (t-1) + b), Wherein, VEM-SYNC (t) is the hidden state, f is the activation function, W_x and W_h are trainable emotion weight matrices, and b is the bias vector. The VEM-SYNC (t) output is used as an emotion weight function and stored in the VEM library.

[0035] 6. The step of the long short-term memory network of the emotion weight function On the basis of the foregoing scheme, the present application includes but is not limited to one or more combinations of the following steps or methods in the step of the long short-term memory network of the emotion weight function: ST610: Establish a long short-term memory network LSTM to realize a long-term memory network model that exceeds the last beat, realize the forward propagation of the emotion, and input the VEM-sync vector emotion value of the forgotten beat or forgotten emotion as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value to fuse and calculate, so as to enhance the emotion representation of the upper and lower beats and obtain the current emotion value.

[0036] ST620: The forgotten beat is an unimportant beat, including but not limited to more than one local accompaniment pass, accompaniment, breath in a song file, or a beat selected by a user, and the forgotten emotion is an unimportant emotion name and its emotion value, or an emotion name and its emotion value selected by a user.

[0037] ST630: The output of the memory network model is an emotion weight function, which is stored in the VEM library.

[0038] 7. Self-attention mechanism step of emotion weight function On the basis of the foregoing scheme, the self-attention mechanism step of the emotion weight function includes but is not limited to one or more combinations of the following model steps or methods: ST710: Obtain the beat number, emotion name and emotion value of the VEM-sync vector of the song file, and mark the VEM-sync vector with the beat number t as VEM-sync(t).

[0039] ST720: Establish a Transformer encoder to realize the forward propagation, backward propagation and omnidirectional propagation of the emotion in the VEM-sync vector sequence of all beats of the song file through the self-attention mechanism, calculate the global dependency relationship between all beat pairs in the VEM-sync vector sequence, and output a group of emotion value sequences enhanced by the global upper and lower VEM-Token sequence weights.

[0040] ST730: The Transformer encoder includes but is not limited to more than one encoder layer, and each encoder layer further includes but is not limited to a multi-head self-attention layer and a feedforward neural network layer, which complete the update and output of the emotion weight function of all VEM-sync vector sequences through residual connection, layer normalization, input embedding and position encoding, wherein the position encoding is obtained from the beat number.

[0041] ST740: Input all VEM-sync vector sequences of the song file to the Transformer encoder, and set one or more emotion names and emotion values that need to be focused on to access the input end of the multi-head self-attention layer, add a multi-head self-attention layer and a feedforward neural network layer in the VEM-sync vector, and store the calculation result of the Transformer encoder in the multi-head self-attention layer.

[0042] ST750: Set a dynamic mechanism selection to access the emotion name and emotion value of the multi-head self-attention layer.

[0043] ST760: Map the emotion name and emotion value of each beat to the query Query, the key Key, and the value Value spaces through a learnable weight matrix, calculate the dot product of the query Query and the key Key on the beat sequence to obtain the correlation weight between the beats, and perform weighted summation on all values Value using the correlation weight to obtain the song file emotion weight information.

[0044] ST770: For the emotion name and emotion value not included in the input of the multi-head self-attention layer, select the top N emotion names, emotion values, and beat numbers with the largest emotion values through the RAG step of setting a prompt word and searching for enhanced generation, and incorporate the emotion weight function to prevent omission, wherein the value of N is determined according to the song file or by the user.

[0045] ST780: Output the result stored in the multi-head self-attention layer as the emotion weight function and store it in the VEM library.

[0046] 8, Enhanced generation RAG step On the basis of the foregoing scheme, the step of outputting the emotion weight function further includes one or more combinations of the following steps or methods: ST810: Obtain the beat number, emotion name, and emotion value of the VEM-sync vector sequence of the song file, and mark the VEM-sync vector with the beat number t as a VEM-sync(t) array, wherein the value range of t includes one or more paragraphs of the song file.

[0047] ST820: Retrieve the VEM library according to the VEM-sync(t) array to obtain one or more VEM-sync vector sequences with the top K-RAG emotion values similar to the local or all emotion names in the VEM-sync(t) array as a reference VEM-sync(t) array sequence, and the similarity is determined by the user.

[0048] ST830: taking the VEM-sync(t) array and the reference VEM-sync(t) array as inputs of the multi-head self-attention mechanism of the Transformer encoder, using the self-attention mechanism to calculate cross-attention weights of the VEM-sync(t) array on the reference VEM-sync(t) array, and then weighting and fusing to obtain a fusion result VEM-sync(t) array.

[0049] ST840: using the fusion result VEM-sync(t) array to verify or correct the preliminary emotion analysis result calculated by the multi-head self-attention network.

[0050] 9. the step of outputting the emotion weight function On the basis of the foregoing scheme, the present application further includes but is not limited to one or more combinations of the following steps or methods in the step of outputting the emotion weight function: ST910: converting the emotion weight function of the entire VEM-sync vector sequence of the song file into an emotion score output in written format, digital format or MIDI communication interface format.

[0051] ST920: establishing a format for communication with the Agent end, and outputting the emotion score to the Agent end according to the emotion weight function of the entire VEM-sync vector sequence of the song file.

[0052] ST930: establishing an NLP-token-oriented interface according to the token standard of the natural language model NLP-token, including a prompt word interface, a context protocol MCP interface and an AI interface.

[0053] ST950: for a song file with video and background pictures that are synchronized in time, establishing a bidirectional picture pointer one that points to and aligns the beat starting point and the picture synchronization point, respectively, according to the tempo and beat starting point of the VEM-Token sequence of the song file.

[0054] ST960: adding a picture pointer two in the VEM-sync vector sequence, which points to and aligns the picture synchronization point according to the tempo and beat starting point of the VEM-Token sequence.

[0055] ST970: establishing a synchronization and alignment relationship between the VEM-sync vector sequence and the video and background pictures through the tempo and beat starting point of the VEM-Token sequence, the synchronization pointer, the picture pointer one and the picture pointer two.

[0056] ST980: According to user needs, output fusion information from VEM-sync vector sequence to video and background picture, including but not limited to song lyrics, song score and information or actions possessed by user needs VEM-sync vector sequence.

[0057] ST990: According to user needs, output fusion information from VEM-sync vector sequence to 3D animation expression action, language action and body action, including but not limited to 3D animation and expression action, language action and body action, to drive 3D animation and expression action, language action and body action, to make expression and body action corresponding to emotion vector in VEM-sync vector sequence.

[0058] ST9A0: Support model context protocol MCP, including MCP Host, MCP Client, MCP Server interface, provide function calling function.

[0059] 10, Intellectual property record step On the basis of the foregoing scheme, the application on the step of video and background picture fusion, specifically includes the following one or more combinations of steps or methods: STA10: Insert intellectual property record mark vector space CV or interface in all VEM-sync vector sequence of song file, used for storing copyright vector.

[0060] STA20: The copyright vector includes but is not limited to copyright holder information, version number, authorization information, encryption and decryption mode.

[0061] STA30: Include encryption and decryption engine, including but not limited to elliptic algorithm, public key, private key mode.

[0062] STA40: Managed by block chain and cloud mode.

[0063] 11, Invention purpose and intention The purpose and intention of the VEM-Token emotion synchronization function hierarchical fusion method of the application are: (1) Realize mapping VEM-Token sequence of song file into VEM-sync emotion synchronization function, and obtain function value of high-dimensional emotion name, emotion value and emotion weight through hierarchical and fusion.

[0064] (2) Cross the transition of discrete text description of NLP-Token, so as to avoid the deviation of discrete natural language text for high-dimensional emotion analog quantity function description.

[0065] (3) Realize dynamic and continuous music emotion, provide guarantee for generating "sound function / emotion function", "sound text / emotion text", "sound animation / emotion animation" and "sound action / emotion action", realize refined recognition and understanding based on vocal music emotion.

[0066] (4) Create automatic and intelligent emotion recognition and understanding of music files, realize multi-modal quantitative representation ability of vocal music emotion.

[0067] (5) Innovate a music / vocal model, access large models to form music AI application agent or dedicated application system, provide strong support for widening AI application.

[0068] 12. Advantages of the invention (1) The VEM-sync emotion synchronization function is adopted to realize the function value of high-dimensional emotion name, emotion value and emotion weight in the VEM-Token sequence describing the song file.

[0069] (2) Avoid the deviation of discrete natural language text for high-dimensional emotion simulation quantity function description.

[0070] (3) Realize dynamic and continuous music emotion, provide guarantee for generating "sound function", "sound text" and "sound animation", realize refined recognition and understanding based on vocal music emotion. BRIEF DESCRIPTION OF DRAWINGS

[0071] LIST OF DRAWINGS Figure 1 : Emotion synchronization function schematic diagram Figure 2 : Emotion vector schematic diagram Figure 3 : Multi-layer scanning synchronization content schematic diagram Figure 4 : Self-attention mechanism and enhanced retrieval generation model schematic diagram Figure 5 : Emotion weight synchronization schematic diagram Figure 6 : Full song emotion synchronization function fusion schematic diagram DETAILED DESCRIPTION OF DRAWINGS The detailed description of each drawing is described in the corresponding drawing number analysis part in the following

specific implementation

[0072] The present application is a patent pool patent of the already granted Chinese invention patent "VEM-Token vocal emotion multi-modal tokenization singing and accompaniment deep learning method, CN120126506" and a published invention patent application "VEM-Token beat capture and alignment model construction method, 202511249168.0". The focus is to invent a new method and new AI model of hierarchical fusion of VEM-Token emotion synchronization function VEM-sync based on synchronization and alignment to the VEM-Token model of the music file.

[0073] The purpose and intention of the present application can be achieved by the following specific embodiments. It needs to be particularly pointed out that each specific embodiment has specific purposes and industrial applicability. Therefore, the following embodiments cannot include all the features and steps of the present application, nor constitute a limitation on the present application. The description of the claims of the present application is the summary of the invention.

[0074] The specific embodiments of the present application are as follows: The innovative VEM-Token emotion synchronization function VEM-sync hierarchical fusion method is a high-precision non-textual music recognition system.

[0075] Diagram explanation The contents of the present embodiment mainly include but are not limited to the following main schematic diagrams, which are: Figures 1 to 6 .

[0076] Implementation step explanation The method steps of the present embodiment mainly include steps 1 to 10. Each of the 10 parts includes several sub-steps. Unless otherwise specified, these sub-steps are not completely required, and unless otherwise specified, the order is not essential, but is optimized and further selected by the patent implementer according to the needs of some specific tasks.

[0077] The specific work steps are as follows: 1. Basic scheme The present application as a VEM-Token emotion synchronization function hierarchical fusion method includes but is not limited to the following steps: ST100: Adopting VEM-Token vocal emotion multi-modal model, dividing the song file into VEM-Token sequence in units of beats, setting the emotion synchronization function as VEM-sync, wherein VEM-sync includes but is not limited to VEM-sync vector sequence corresponding to and aligned with VEM-Token sequence, VEM-sync vector includes but is not limited to synchronization content and synchronization pointer pointing to corresponding VEM-Token beat.

[0078] ST200: According to the VEM-Token sequence, the synchronization content of the VEM-sync vector sequence is calculated, including but not limited to beat attributes and emotion attributes, wherein the emotion attributes include but are not limited to one or more of emotion names, emotion values, and emotion weights.

[0079] ST300: The emotion weight function includes but is not limited to: using hierarchical processing of song files to obtain corresponding VEM-sync vector sequences, and hierarchical emotion weights for emotion values therein, using forward propagation, backward propagation, and omnidirectional propagation including but not limited to multi-layer weighted scanning, recurrent neural networks, long short-term memory networks, and self-attention mechanisms, and fusing to obtain emotion synchronization functions with song files.

[0080] Among them, the VEM-Token vocal emotion multi-modal model refers to the model in "VEM-Token vocal emotion multi-modal tokenization song and accompaniment deep learning method, CN120126506", which specifically includes steps 1-4: (1) Record emotions using one or more modalities, label the vocal emotion multi-modal as VEM, and construct VEM classification, VEM coordinate system, VEM function, and VEM library. Vocal emotion includes one or a combination of joy, sadness, anger, fear, disgust, surprise, calm, anticipation, trust, love, hate, emotion, and enmity. Multi-modal includes one or a combination of lyrics, song, accompaniment, vocal style, music, emotion base, accompaniment instrument, video, and image. The VEM coordinate system includes a coordinate axis system established according to independent emotions, opposite emotion pairs, and associated opposite emotion groups.

[0081] (2) Collect vocal samples according to VEM classification, and perform emotion evaluation on vocal samples in song and accompaniment by human vocal experts. Use supervised learning and deep learning to train VEM functions to obtain VEM parameters and add them to the VEM library.

[0082] (3) Use VEM processor to perform beat calibration on vocal files, and separate song stream and accompaniment stream. According to the beat, perform VEM-Token segmentation on the vocal file, convert the song stream into a VEM-Token1 sequence, and convert the accompaniment stream into a VEM-Token2 sequence, and add them to the preprocessing library.

[0083] (4) Use deep learning to generate lyrics score, VEM-Token song score, VEM-Token accompaniment score, and VEM-Token score.

[0084] The model of VEM-Token beat capture and beat alignment refers to the model in the VEM-Token beat capture and alignment model construction method (202511249168.0), and specifically includes steps 5-7. (5) For a vocal file, a beat model including beat capture and beat alignment is set according to a VEM-Token vocal emotion multi-modal model, the beats of the vocal file are captured, and the vocal file is divided into VEM-Token sequences according to the beats, and the positions of the start of the beat and the end of the beat in each VEM-Token are marked.

[0085] (6) A start point alignment model is set, including: The sample file included in the vocal file and the user file sung by the user imitating the sample file are divided into VEM-Token1 sequences and VEM-Token2 sequences, respectively, and the start points of each VEM-Token1 are used to adjust the start points of the corresponding VEM-Token2 one by one to align with the start points of the corresponding VEM-Token1.

[0086] For each segment of the loop segment included in the vocal file, the start points of each VEM-Token of each segment are adjusted one by one using the start point fine-tuning step to align with the start points of the VEM-Tokens at the corresponding positions of the first segment, starting from the second segment, until all loop segments end.

[0087] (7) A terminal point alignment model is set, including: According to the terminal points of each VEM-Token1, the terminal points of the corresponding VEM-Token2 are adjusted one by one using the terminal point fine-tuning step to align with the terminal points of the corresponding VEM-Token1.

[0088] For each segment of the loop segment, the terminal points of each VEM-Token of each segment are adjusted one by one using the terminal point fine-tuning step to align with the terminal points of the VEM-Tokens at the corresponding positions of the first segment, starting from the second segment, until all loop segments end.

[0089] In the present application, although the construction of the emotion synchronization function or model is based on the ideas and inspirations of the two patents, it is also an independent invention, which is to further realize the structured analysis of the structured system, order and post-structured after the realization of the structured model of the emotion summary of the song file.

[0090] As Figure 1 This is a schematic diagram of the emotion synchronization function of the present application, which is analyzed as follows: Figure 1The lower part is the VEM-Token sequence divided by the song file, and the file attribute is a spectrum format file. The upper part is the emotion synchronization function VEM-sync, Beat is the beat, and n is the number of the sequence. The design intends to: (1) VEM-Token is a division of the song file in a certain structure, including the VEM-Token vocal emotion multi-modal tokenization song and accompaniment deep learning method CN120126506, the VEM-Token vocal emotion multi-modal division model of the method for constructing a beat capture and alignment model 202511249168.0, and other division methods such as MIDI and MCS models. On the VEM-Token sequence of the song file, a VEM-sync vector sequence is attached, wherein the VEM-sync vector sequence corresponds to the VEM-Token sequence one by one.

[0091] (2) The correspondence between the VEM-sync vector sequence and the VEM-Token sequence is constrained by the synchronization pointer. It should be noted that although Figure 1 the VEM-sync vector sequence and the VEM-Token sequence in the above-mentioned CN120126506 are "equal length", this equal length is only logical and not equal length in the physical array sense, because the emotion name, emotion value and emotion weight parsed out are different with different emotion complexity of the VEM-Token sequence of the song, and the storage space size occupied is also different.

[0092] (3) Since the VEM-Token sequence is divided according to the music beat, and the beat is the most basic unit of the music language, the present application takes the beat unit as the beat layer, and further divides the bar layer, the sentence layer and the whole song layer on the beat layer. According to this layer division, the VEM-sync vector sequence is correspondingly synchronized into the beat layer, the bar layer, the sentence layer and the whole song layer. It should be noted that this physical and time sequence division method according to the bar layer, the sentence layer and the whole song layer is only one of the division levels of the present application, and is not unique. The present application also includes emotion, logic and spatial division methods.

[0093] (4) The emotion synchronization function VEM-sync, the so-called "synchronization" means that the VEM-sync vector corresponds to the corresponding VEM-Token according to the synchronization pointer; the so-called "emotion synchronization" means that the emotion description in the VEM-sync vector is synchronized with the VEM-Token; and the so-called "layered fusion" of the present application means that the emotion analysis result of the whole song is obtained according to the "layered" analysis and "fusion" according to the layered analysis.

[0094] (5) It is noted that the VEM-sync vector is a quantity with magnitude and direction, which can be represented as a matrix or an array in mathematics. Its spatial direction information has been incorporated into the data structure. Therefore, in the present invention, the VEM-sync vector, vector and array are considered as the same concept and are not distinguished.

[0095] 2. VEM-Token vocal emotion multi-modal model On the basis of the foregoing basic scheme, the present invention in the VEM-Token vocal emotion multi-modal model includes but is not limited to one or a combination of the following beat capture and alignment steps or methods: ST110: The time measurement of the VEM-Token sequence is in seconds or beat numbers, with a length starting from 0 to the end of the song.

[0096] ST120: The synchronization pointer includes a single synchronization pointer pointing to and aligning with the VEM-Token beat start time, or a double synchronization pointer pointing to the VEM-Token beat start and end times respectively.

[0097] ST130: The length of the data structure of the synchronization content includes but is not limited to fixed length and variable length.

[0098] ST140: The VEM classification refers to the classification of emotions, including but not limited to independent emotions, opposite emotion pairs, and related opposite emotion groups, which in turn establish one or a combination of one-dimensional unidirectional, one-dimensional bidirectional, and multi-dimensional bidirectional VEM coordinate systems.

[0099] ST150: The VEM modalities include but are not limited to one or a combination of lyrics, singing, accompaniment, style, music, accompaniment instruments, video, and image.

[0100] ST160: The style includes but is not limited to one or a combination of national song singing, popular song singing, Western song singing, pop song singing, original song singing, and opera singing.

[0101] ST170: The emotion name is one or a combination of joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, emotion, and hatred included in the VEM modalities, and the emotion value is set to the emotion measurement including but not limited to percentage or numerical value.

[0102] ST180: The VEM library includes but is not limited to at least VEM modalities, styles, emotion names, and emotion values, which are generated by supervised learning, reinforcement learning by human vocal experts for typical song guidance, and deep learning by the VEM-Token vocal emotion multi-modal model.

[0103] In the present application, the emotion is expressed and modeled by a vector with a length and a direction, and an emotion is characterized by an emotion name and an emotion value. There are also relationships between emotions, including at least independent emotions, opposite emotion pairs, and related opposite emotion groups, which in turn establish a VEM coordinate system including but not limited to one or a combination of one-dimensional unidirectional, one-dimensional bidirectional, and multi-dimensional bidirectional.

[0104] As shown in Figure 2 , this is a schematic diagram of the emotion vector, which is analyzed as follows: (1) For related opposite emotion groups, such as love, hate, emotion, enemy, joy, and sorrow, here, love, hate, emotion, enemy, joy, and sorrow are 3 groups of opposite emotions, and there is a certain correlation between the emotion groups. In order to facilitate mathematical analysis, the 3 groups of emotions are included in a three-dimensional coordinate system in an orthogonal manner, Figure 2 is this example. It should be noted that according to the theory of psychology, there are other modeling methods between emotion groups, and this application only shows one modeling method of emotion groups. Potential users of this patent application can use other modeling methods.

[0105] (2) Figure 2 , the emotion value of a certain E point of the 3 emotion groups of love, hate, emotion, enemy, joy, and sorrow is marked by the three-dimensional coordinate system such as the three-dimensional orthogonal coordinate system XYZ E(x, y, z). It should be noted that since the 3 groups of emotions are related to each other (orthogonal relationship), the E point on the X, Y, and Z coordinate axes produces projection values: x, y, and z, and also has geometric and psychological significance. For establishing a mathematical model, it also has practical value for subsequent calculation.

[0106] It should be noted that the VEM library is based on the existing VEM library before the present application is used. Its early stage is completed by a human vocal music expert, specifically by collecting some typical songs according to the vocal music style and VEM modal classification, and then generating the results after supervised learning and reinforcement learning of artificial intelligence. The later stage is to perform VEM-Token vocal emotion multi-modal model by artificial intelligence, and then merge the results with the previous results after deep learning.

[0107] 3. Basic steps of synchronization content On the basis of the foregoing scheme, the present application includes but is not limited to one or more of the following steps or methods in the basic steps of synchronization content: ST210: The beat attribute is calculated by the VEM-Token vocal music emotion multi-modal model or obtained from the score of the song file. The beat attribute specifically includes but is not limited to: a measure is composed of more than one beat, the time unit uses the number of beats per minute, the beat type includes but is not limited to strong, secondary strong, weak, and rest, and the type includes but is not limited to 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, 3 / 8 beat, 6 / 8 beat, and 7 / 8 beat.

[0108] ST220: The emotion attribute is calculated by the VEM-Token vocal music emotion multi-modal model for the VEM-Token sequence and more than one emotion name. The emotion value of the emotion name is calculated and normalized to between -100% and +100%. The emotion weight is set, and the emotion name, emotion value, and emotion weight are combined to form an emotion weight function storage unit and stored in the VEM-sync vector sequence.

[0109] ST230: The emotion weight function is defined as a weighting function for an emotion value to modify the resulting emotion value of an emotion name in different layers in a VEM-sync vector in the VEM-sync vector sequence.

[0110] ST240: The layer is set as the VEM-sync vector sequence of the song file, which is divided into the beat layer, measure layer, sentence layer, and whole song layer, and the emotion weight function is applied layer by layer to obtain the corresponding weighted emotion value.

[0111] ST250: According to the emotion change of the entire song file, the emotion weight function is used for hierarchical fusion to calculate and obtain the emotion synchronization function of the entire song file.

[0112] It should be noted that, in addition to the vertical division method from the bottom (i.e., the beat layer) to the top (i.e., the measure layer, sentence layer, and whole song layer in turn), there is also a horizontal division method according to, for example, an emotion name or an instrument name. The emotion attribute transmission includes vertical transmission and horizontal transmission, and the horizontal transmission can be further divided into forward transmission, backward transmission, and bidirectional transmission.

[0113] 4. Multi-layer weighted scanning step On the basis of the foregoing scheme, the multi-layer weighted scanning step of the emotion weight function of the present application includes but is not limited to one or more combinations of the following steps or methods: ST410: The beat attribute is calculated by the VEM-Token vocal music emotion multi-modal model or obtained from the score of the song file. The beat attribute specifically includes but is not limited to: a measure is composed of more than one beat, the time unit uses the number of beats per minute, the beat type includes but is not limited to strong, secondary strong, weak, and rest, and the type includes but is not limited to 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, 3 / 8 beat, 6 / 8 beat, and 7 / 8 beat.

[0114] ST420: Add the phrase layer to the VEM-sync vector sequence in units of phrases, average the emotion values of the same emotion name in the phrase layer, and store each category of emotion and the corresponding average emotion value in the phrase layer.

[0115] ST430: Add the sentence layer to the VEM-sync vector sequence in units of sentences, accumulate the emotion values of the same emotion name in the sentence layer, extract the style weight from the VEM library according to the style, or select the style weight from the user, obtain the emotion values of the emotion name after weighted calculation by using the weighted algorithm, and store each category of emotion and the corresponding weighted average emotion value in the sentence layer.

[0116] ST440: Add the whole song layer to the VEM-sync vector sequence in units of whole songs, sort the emotion values of the emotion name in the sentence layer from large to small, select the emotion name of the first part or the emotion name specified by the user as the high-weight emotion, wherein the first part is less than 40% according to the statistical results in the VEM library or selected by the user, store the high-weight emotion and the corresponding emotion value in the whole song layer, output as the emotion weight function, and store in the VEM library.

[0117] As Figure 3 , this is a bottom-up multi-layer scanning synchronization content diagram, which is analyzed as follows: (1) Figure 3 In the above, a 4 / 4 song is divided into four levels from bottom to top, namely the beat layer, the phrase layer, the sentence layer and the whole song layer. Among them, the VEM-Token vocal emotion multi-modal model is used to first calculate the division of the beat, and to calculate all the emotion names and emotion values of the emotion name, which are included in the corresponding VEM-sync vector sequence.

[0118] (2) Since the song is a 4 / 4 song, the emotion name and emotion value of every four beats are included in a phrase layer. After the emotion name is summarized, the emotion value is calculated and stored in the phrase layer.

[0119] (3) The so-called sentence layer refers to the sentences in the song divided according to the lyrics, and each sentence includes several phrases. The emotion value included in the sentence layer is obtained by querying in the VEM library according to the style, or selected by the user.

[0120] (4) All the sentence layers are fused into the emotion sequence of the whole song. It should be noted that in the whole song layer, the corresponding emotion sequence is usually calculated according to the order of the sentence layer, that is, the emotion of the whole song layer is not a single value, but a value that changes with the beat time flow according to the emotion of the song. Therefore, the present application refers to it as an emotion function.

[0121] 5. The step of recurrent neural network On the basis of the foregoing scheme, the step of the emotion weight function of the recurrent neural network of the present application includes but is not limited to one or more combinations of the following steps or methods: ST510: Establish a recurrent neural network RNN to realize a memory network model, realize the forward propagation of emotion, and input the emotion value of the VEM-sync vector of the previous beat as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value to fuse and calculate, so as to enhance the emotion representation of the upper and lower beats, and obtain the current emotion value.

[0122] ST520: Obtain the emotion name and emotion value of the VEM-sync vector, mark the VEM-sync vector with beat sequence number t as the hidden history state VEM-sync(t) of the current beat, mark the VEM-sync vector with beat sequence number t-1 as the hidden history state VEM-sync(t-1) of the previous beat, and t starts from 2 until the end of the song.

[0123] ST530: Set the calculation formula of the memory network model as: VEM-SYNC(t)=f(W_x▪VEM-sync(t)+W_h▪VEM-sync(t-1)+b), Wherein, VEM-SYNC(t) is the hidden state, f is the activation function, W_x and W_h are trainable emotion weight matrices, and b is the bias vector. The VEM-SYNC(t) output is stored in the VEM library as an emotion weight function.

[0124] This is a typical scheme of emotion forward propagation based on recurrent neural network RNN. By modifying the recurrent network, a scheme of vertical propagation of convolutional CNN can be supported.

[0125] 6. The step of long short-term memory network On the basis of the foregoing scheme, the step of the emotion weight function of the long short-term memory network of the present application includes but is not limited to one or more combinations of the following steps or methods: ST610: Establish a long short-term memory network LSTM to realize a long-term memory network model exceeding the previous beat, realize the forward propagation of emotion, and input the emotion value of the VEM-sync vector of the forgotten beat or forgotten emotion as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value to fuse and calculate, so as to enhance the emotion representation of the upper and lower beats, and obtain the current emotion value.

[0126] ST620: The forgotten beat is the beat that is not important, including but not limited to more than one local accompaniment, door, singing, breath in the song file, or the beat selected by the user, and the forgotten mood also includes the mood name and its mood value that is not important, or the mood name and its mood value selected by the user.

[0127] ST630: The output of the memory network model is the mood weight function, which is stored in the VEM library.

[0128] It should be noted that in the VEM library, there are contents previously annotated by human vocal music experts, including the mark of the memory network model, especially for some special songs, the audience's mood sound in the singing and performance environment is also recorded and marked.

[0129] 7. Self-attention mechanism step On the basis of the foregoing scheme, the self-attention mechanism step of the mood weight function includes but is not limited to one or more combinations of the following model steps or methods: ST710: Obtain the mood name and mood value of the VEM-sync vector of the song file, and mark the VEM-sync vector with the beat sequence number t as VEM-sync(t).

[0130] ST720: Establish a Transformer encoder to realize the forward propagation, backward propagation and omnidirectional propagation of the mood in the VEM-sync vector sequence of all beats of the song file through the self-attention mechanism, calculate the global dependency relationship between all beat pairs in the VEM-sync vector sequence, and output a group of mood value sequences enhanced by global up and down VEM-Token sequence weights.

[0131] ST730: The Transformer encoder includes but is not limited to more than one encoder layer, and each encoder layer further includes but is not limited to a multi-head self-attention layer and a feedforward neural network layer, which updates and outputs the mood weight function for all VEM-sync vector sequences through residual connection, layer normalization, input embedding and position encoding, wherein the position encoding is obtained from the beat sequence number.

[0132] ST740: Input all VEM-sync vector sequences of the song file to the Transformer encoder at one time, and set one or more mood names and mood values that need to be focused on to access the input end of the multi-head self-attention layer, wherein the input end is v, and the VEM-sync vector is added to include but not limited to the multi-head self-attention layer and the feedforward neural network layer, and the Transformer encoder stores the calculation result in the multi-head self-attention layer.

[0133] ST750: Set the dynamic mechanism selection to access the multi-head self-attention layer of the emotion name and emotion value.

[0134] ST760: Map the emotion name and emotion value of each beat to the query, key, and value spaces through a learnable weight matrix. Calculate the dot product of the query and key on the beat sequence to obtain the association weight between beats. Then, perform a weighted sum of all values using this association weight to obtain the song file emotion weight information.

[0135] ST770: For emotion names and emotion values not included in the multi-head self-attention layer input, set the prompt word and use the RAG step to generate the emotion value. Select the top N emotion names, emotion values, and beat numbers with the largest emotion value and incorporate them into the emotion weight function to prevent omission. The value of N is determined according to the song file or by the user.

[0136] ST780: Output the result stored in the multi-head self-attention layer as the emotion weight function and store it in the VEM library.

[0137] As shown in Figure 4 , this is a schematic diagram of the self-attention mechanism step of the emotion weight function. The analysis is as follows: Figure 4 It includes three parts: the emotion synchronization function VEM-sync, the VEM-Token sequence, and the video background. The VEM-Token sequence is obtained by using the VEM-Token vocal emotion multi-modal model after beat division and alignment. It is worth noting that the VEM-Token sequence here is actually a spectral format file with beat information added. The video background is the MV (Music Video) in the song file, including moving pictures and still pictures.

[0138] (2) The emotion synchronization function VEM-sync is a fixed-length or variable-length array, also known as a matrix. The synchronization pointers point to the VEM-Token sequence and the beat part of the video background divided according to the beat, respectively; the VEM-sync vector includes a two-dimensional array composed of multiple emotion names, emotion values, and emotion weights; the emotion name and emotion value are calculated by the VEM-Token vocal emotion multi-modal model for the corresponding VEM-Token sequence.

[0139] (3) The emotion synchronization function VEM-sync here includes a Transformer encoder structure and works with a multi-head self-attention mechanism, and also includes a RAG structure and works with an enhanced retrieval generation mechanism.

[0140] ​(4) It should be noted that the division of labor or the adoption or not of the multi-head self-attention layer and the RAG layer can be determined by the user according to the attributes of the emotion.

[0141] Figure 5 is an emotion weight synchronization schematic diagram. It not only has a reference value for the schematic reference of the present step, but also has a reference value for the aforementioned multi-layer weighted scanning step. In Figure 5 , it can be divided into the left emotion attribute initial matrix including the emotion name, emotion value and emotion weight, the middle emotion synchronization function, and the right emotion attribute result matrix including the emotion name, emotion value and emotion weight. Here, the emotion synchronization function calculates the emotion weight in the initial matrix through the present step and stores the result in the result matrix.

[0142] 8. Enhanced RAG step On the basis of the foregoing scheme, the present application further includes but is not limited to one or more combinations of the following steps or methods in the output step of the emotion weight function: ST810: Obtain the tempo, emotion name and emotion value of the VEM-sync vector sequence of the song file, and mark the VEM-sync vector with the tempo t as the VEM-sync(t) array, wherein the value range of t includes one or more paragraphs of the song file.

[0143] ST820: Retrieve the VEM library according to the VEM-sync(t) array, obtain one or more VEM-sync vector sequences of the first K-RAG emotion values similar to the local emotion name or all emotion names in the VEM-sync(t) array as the reference VEM-sync(t) array sequence, and the similarity is determined by the user.

[0144] ST830: Take the VEM-sync(t) array and the reference VEM-sync(t) array as the input end of the multi-head self-attention mechanism of the Transformer encoder, adopt the self-attention mechanism, perform weighted fusion on the VEM-sync(t) array after calculating the cross-attention weight on the reference VEM-sync(t) array, and obtain the fusion result VEM-sync(t) array.

[0145] ST840: The fusion result VEM-sync(t) array is used to verify or correct the preliminary emotion analysis result calculated by the multi-head self-attention network.

[0146] See Figure 5 and the description.

[0147] It should be noted that in the step ST830, "the VEM-sync (t) array is weighted and fused after the cross-attention weight is calculated with the reference VEM-sync (t) array", which includes vertical merging or horizontal merging of the two groups of matrices.

[0148] 9. The step of outputting the emotion weight function Based on the foregoing scheme, the present application further includes but is not limited to one or more combinations of the following steps or methods in the step of outputting the emotion weight function: ST910: According to the emotion weight function of the entire VEM-sync vector sequence of the song file, the emotion score output is converted into a written format, a digital format or a MIDI communication interface format.

[0149] ST920: Establish a format for communication with the Agent end, and output the emotion score to the Agent end according to the emotion weight function of the entire VEM-sync vector sequence of the song file.

[0150] ST930: According to the standard of NLP-token word element of natural language model, establish an NLP-token interface, including a prompt word interface, a context protocol MCP interface and an AI interface.

[0151] ST950: For a song file with video and background pictures synchronized in time, according to the time signature and beat starting point of the VEM-Token sequence of the song file, a bidirectional picture pointer one is established, which points to and aligns the beat starting point and the picture synchronization point respectively.

[0152] ST960: Add a picture pointer two in the VEM-sync vector sequence, which points to and aligns the picture synchronization point according to the time signature and beat starting point of the VEM-Token sequence.

[0153] ST970: Through the time signature and beat starting point of the VEM-Token sequence, the synchronization pointer, the picture pointer one and the picture pointer two, the synchronization and alignment relationship between the VEM-sync vector sequence and the video and background pictures is established.

[0154] ST980: According to the user's demand, the fusion information is output from the VEM-sync vector sequence to the video and background pictures, including but not limited to the lyrics score, the song score and the information or action possessed in the VEM-sync vector sequence according to the user's demand.

[0155] ST990: According to user needs, the fusion information outputted by the VEM-sync vector sequence to include but not limited to 3D animation expression action, language action and body action, to drive the 3D animation and expression action, language action and body action, to make the expression and body action corresponding to the emotion vector in the VEM-sync vector sequence.

[0156] ST9A0: Support model context protocol MCP (Model Context Protocol), including MCP Host, MCPClient, MCP Server interface, provide function calling function Function Calling. In order to use the application with other large models or AI systems.

[0157] Figure 6 is a full song emotion synchronization function fusion schematic diagram. The analysis is as follows: (1) Figure 6 In it, 601 is the emotion projection slice of the node in the full song layer, 602 is the emotion synchronization function vector, and 603 is the beat line.

[0158] (2) The emotion projection slice is a diagram of the emotion synchronization function at a moment. Since the emotion synchronization function is the instantaneous projection of the multi-dimensional emotion vector on the beat time axis, in order to facilitate explanation, this emotion projection slice is drawn as one containing only love, hate, emotion, and hatred. In fact, it should be the instantaneous projection of the synthesis or fusion of all emotion categories.

[0159] (3) For the full song, the emotion synchronization function vector is a kind of high-dimensional continuous variable according to the full song time.

[0160] 10. The step of intellectual property record On the basis of the foregoing scheme, the application further includes but is not limited to the steps of intellectual property record and management, specifically including the following one or more combined steps or methods: STA10: Insert the mark vector space CV or interface of the intellectual property record in all VEM-sync vector sequences of the song file, used to store the copyright vector.

[0161] STA20: The copyright vector includes but is not limited to copyright holder information, version number, authorization information, encryption and decryption mode.

[0162] STA30: Including encryption and decryption engine, including but not limited to elliptic algorithm, public key, private key mode.

[0163] STA40: Managed by block chain and cloud mode.

[0164] It should be noted that one of the features that distinguishes the present application from other methods is the VEM-sync vector sequence with the emotion synchronization function, which provides a storage space outside the song file and its VEM-Token sequence, so that the storage of intellectual property rights using the storage space of the VEM-sync vector sequence does not affect the song file itself and is also conducive to the storage of encryption and decryption.

Claims

1. A method for hierarchical fusion of VEM-Token emotion synchronization functions, characterized in that, Comprising: ST100: A VEM-Token vocal emotion multi-modal model is adopted to segment a song file into a VEM-Token sequence in a beat unit, and a emotion synchronization function VEM-sync is set, wherein the VEM-sync includes a VEM-sync vector sequence corresponding to and aligned with the VEM-Token sequence, and the VEM-sync vector includes synchronization content and a synchronization pointer pointing to a corresponding VEM-Token beat; ST200: According to the VEM-Token sequence, the synchronization content of the VEM-sync vector sequence is calculated, and the synchronization content includes beat attributes and emotion attributes, wherein the emotion attributes include one or more emotion names, emotion values, and emotion weights; ST300: The emotion weight function includes: adopting hierarchical processing of a song file to obtain a corresponding VEM-sync vector sequence, and adopting one or a combination of multi-layer weighted scanning, recurrent neural network, long short-term memory network, self-attention mechanism, and retrieval enhancement to perform forward propagation, backward propagation, and omnidirectional propagation to obtain an emotion synchronization function of the song file.

2. The method according to claim 1, characterized in that, The VEM-Token vocal emotion multi-modal model comprises: ST110: The time measurement of the VEM-Token sequence is seconds or beats, and the length starts from 0 to the end of the song; ST120: The synchronization pointer includes: a single synchronization pointer pointing to and aligned with the start time of the VEM-Token beat, or a double synchronization pointer pointing to the start and end times of the VEM-Token beat, respectively; ST130: The data structure of the synchronization content includes fixed length and variable length; ST140: The VEM classification refers to the classification of emotions, including independent emotions, opposite emotion pairs, and associated opposite emotion groups, and a VEM coordinate system is established in one or a combination of one-dimensional unidirectional, one-dimensional bidirectional, and multi-dimensional bidirectional; ST150: The VEM modalities include one or a combination of lyrics, singing, accompaniment, style, music, accompaniment instruments, video, and image; ST160: The style includes one or a combination of national song singing method, popular song singing method, western song singing method, pop song singing method, original song singing method, and opera singing method; ST170: The emotion name is one or a combination of joy, sadness, anger, fear, disgust, surprise, calm, expectation, trust, love, hate, emotion, and hatred in the VEM modalities, and the emotion value is set to be a percentage or a numerical value for emotion measurement; ST180: The VEM library includes at least VEM modalities, styles, emotion names, and emotion values, which are generated by a VEM-sync vector sequence supervised learning, reinforcement learning, and deep learning performed by a VEM-Token vocal emotion multi-modal model under the guidance of a typical song by a human vocal expert.

3. The method according to claim 2, characterized in that, The basic steps of the synchronization content specifically include: ST210: The beat attribute is calculated by the VEM-Token vocal music emotion multi-modal model or obtained from the music file score collection. The beat attribute specifically includes: time signature, measure, beat type, and type, wherein the measure is composed of more than one beat, the time unit uses the number of beats per minute, the beat type includes strong, secondary strong, weak, and rest, and the type includes 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, 3 / 8 beat, 6 / 8 beat, and 7 / 8 beat; ST220: The emotion attribute is calculated by the VEM-Token vocal music emotion multi-modal model for the VEM-Token sequence and more than one emotion name, and the emotion value of the emotion name is normalized to between -100% and +100%. The emotion weight is set, and the time signature, emotion name, emotion value, and emotion weight are combined to form an emotion weight function, which is stored in the VEM-sync vector sequence; ST230: The emotion weight function is defined as a weighting function for an emotion value to modify the resulting emotion value of an emotion name in different layers in a VEM-sync vector in the VEM-sync vector sequence; ST240: The layer is set as the VEM-sync vector sequence of the music file, which is divided into beat layer, measure layer, sentence layer, and whole song layer, and the emotion weight function is applied layer by layer to obtain the corresponding weighted emotion value; ST250: According to the emotion change of the whole music file, the emotion weight function is calculated by the hierarchical fusion method to obtain the emotion synchronization function of the whole music file.

4. The method according to claim 3, characterized in that, The emotion weight function includes a multi-layer weighting scanning step, specifically including: ST410: The beat unit is used to obtain the beat layer emotion attribute of the music file; ST420: The measure unit is used to add the measure layer in the VEM-sync vector sequence, the average value of the emotion values of the same emotion name in the measure layer is taken, and the emotion of each type and the corresponding emotion average value are stored in the measure layer; ST430: The sentence unit is used to add the sentence layer in the VEM-sync vector sequence, the emotion values of the same emotion name in the sentence layer are accumulated, the style weight is extracted from the VEM library according to the style, or the style weight is selected by the user, the weighted emotion value of the emotion name is obtained by using the weighting algorithm, and the emotion of each type and the corresponding weighted emotion average value are stored in the sentence layer; ST440: The whole song unit is used to add the whole song layer in the VEM-sync vector sequence, the emotion values of the emotion names in the sentence layer are sorted from large to small, the first part of the emotion names or the user-specified emotion names are selected as high-weight emotions, wherein the first part is less than 40% according to the statistical result in the VEM library or selected by the user, the high-weight emotions and the corresponding emotion values are stored in the whole song layer, and the emotion weight function is output and stored in the VEM library.

5. The method according to claim 3, characterized in that, The emotion weight function includes a recurrent neural network step, specifically including: ST510: Establish a recurrent neural network (RNN) to implement a memory network model, implement forward propagation of emotions, and input the emotion value of the VEM-sync vector of the previous beat as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value to calculate the fusion, to enhance the emotion representation of the upper and lower beats, and obtain the current emotion value; ST520: Obtain the emotion name and emotion value of the VEM-sync vector, mark the VEM-sync vector with beat number t as the hidden history state VEM-sync(t) of the current beat, and mark the VEM-sync vector with beat number t-1 as the hidden history state VEM-sync(t-1) of the previous beat, t starts from 2 until the end of the song; ST530: Set the calculation formula of the memory network model as: VEM-SYNC(t) = f(W_x ▪ VEM-sync(t) + W_h ▪ VEM-sync(t-1) + b), where VEM-SYNC(t) is the hidden state, f is the activation function, W_x and W_h are trainable emotion weight matrices, and b is the bias vector. The VEM-SYNC(t) output is stored as an emotion weight function in the VEM library.

6. The method according to claim 3, characterized by The emotion weight function includes the steps of a long short-term memory network, specifically including: ST610: Establish a long short-term memory network (LSTM) to implement a longer-term memory network model than the previous beat, implement forward propagation of emotions, and input the emotion value of the VEM-sync vector of the forgotten beat or forgotten emotion as the hidden history state of the current beat when calculating the VEM-sync vector of the current beat, combine the current emotion value to calculate the fusion, to enhance the emotion representation of the upper and lower beats, and obtain the current emotion value; ST620: The forgotten beat is an unimportant beat, including one or more local accompaniment, pass door, singing, breath in the song file, or the beat selected by the user, and the forgotten emotion is an unimportant emotion name and its emotion value, or the emotion name and its emotion value selected by the user; ST630: The output of this memory network model is the emotion weight function, which is stored in the VEM library.

7. The method according to claim 3, characterized in that, The emotion weight function includes the steps of a self-attention mechanism, specifically including: ST710: Obtain the beat number, emotion name, and emotion value of the VEM-sync vector of the song file, and mark the VEM-sync vector with beat number t as VEM-sync(t); ST720: Establish a Transformer encoder to implement forward propagation, backward propagation, and omnidirectional propagation of emotions in the VEM-sync vector sequence of all beats of the song file through a self-attention mechanism, calculate the global dependency relationship between all beat pairs in the VEM-sync vector sequence, and output a set of emotion value sequences enhanced by global upper and lower VEM-Token sequence weights; ST730: The Transformer encoder includes one or more encoder layers, each of which further includes a multi-head self-attention layer and a feed-forward neural network layer, which updates and outputs the emotion weight function for the entire VEM-sync vector sequence through a residual connection, layer normalization, input embedding, and position encoding, where the position encoding is taken from the beat number; ST740: The Transformer encoder is inputted with the entire VEM-sync vector sequence of the song file, and one or more emotion names and emotion values that need to be focused on are set to access the input end of the multi-head self-attention layer, and the multi-head self-attention layer and the feed-forward neural network layer are added in the VEM-sync vector, and the Transformer encoder stores the calculation result in the multi-head self-attention layer; ST750: Set up a dynamic mechanism selection to access the emotion name and emotion value of the multi-head self-attention layer; ST760: Map each beat's emotion name and emotion value to the Query, Key, and Value spaces through a learnable weight matrix, calculate the dot product of Query and Key on the beat sequence to obtain the association weight between beats, and use this association weight to weight sum all values Value to obtain the song file emotion weight information; ST770: For emotion names and emotion values not included in the multi-head self-attention layer input, select the top N emotion names, emotion values, and beat numbers with the largest emotion values by setting a prompt word and going through the RAG step of retrieval-enhanced generation, and incorporate them into the emotion weight function to prevent omissions, where the value of N is determined according to the song file or by the user; ST780: Output the result stored in the multi-head self-attention layer as the emotion weight function and store it in the VEM library.

8. The method according to claim 7, characterized in that, The emotion weight function includes the RAG step of retrieval-enhanced generation, which specifically includes: ST810: Obtain the beat number, emotion name, and emotion value of the VEM-sync vector sequence of the song file, and mark the VEM-sync vector with beat number t as the VEM-sync(t) array, where the value of t includes one or more paragraphs of the song file; ST820: Retrieve the VEM library according to the VEM-sync(t) array to obtain one or more VEM-sync vector sequences with the top K-RAG emotion values similar to the local or all emotion names in the VEM-sync(t) array as reference VEM-sync(t) array sequences, where the similarity is determined by the user; ST830: Take the VEM-sync(t) array and the reference VEM-sync(t) array as the input end of the multi-head self-attention mechanism of the Transformer encoder, and use the self-attention mechanism to calculate the cross-attention weight between the VEM-sync(t) array and the reference VEM-sync(t) array, and then perform weighted fusion to obtain the fusion result VEM-sync(t) array; ST840: The fusion result VEM-sync (t) array is used to verify or correct the preliminary emotion analysis result calculated by the multi-head self-attention network.

9. The method of any of claims 4, 5, 6, 7, 8, wherein, The emotion weight function includes the following steps: ST910: According to the emotion weight function of the entire VEM-sync vector sequence of the song file, convert the output into a written format, a digital format, or a MIDI communication interface format of the emotional score; or, ST920: Establish a format for communication with the Agent end, and output the emotional score to the Agent end according to the emotion weight function of the entire VEM-sync vector sequence of the song file; Or, ST930: According to the standard of NLP-token word units of the natural language model, establish an NLP-token interface, including a prompt word interface, a context protocol MCP interface, and an AI interface; or, ST950: For a song file with a video and a background picture that are synchronized in time, according to the tempo and beat start point of the VEM-Token sequence of the song file, establish a bidirectional picture pointer one that points to and aligns the beat start point and the picture synchronization point, respectively; ST960: Add a picture pointer two in the VEM-sync vector sequence, which points to and aligns the picture synchronization point according to the tempo and beat start point of the VEM-Token sequence; ST970: Establish a synchronization and alignment relationship between the VEM-sync vector sequence and the video and background picture through the tempo and beat start point of the VEM-Token sequence, the synchronization pointer, the picture pointer one, and the picture pointer two; ST980: According to user requirements, output fusion information from the VEM-sync vector sequence to the video and background picture, including lyrics, song scores, and information or actions in the VEM-sync vector sequence that meet user requirements; ST990: According to user requirements, output fusion information from the VEM-sync vector sequence to include 3D animation expression actions, language actions, and body actions to drive 3D animation and expression actions, language actions, and body actions to make expressions and body actions corresponding to the emotion vectors in the VEM-sync vector sequence; ST9A0: Support the model context protocol MCP, including the MCP Host, MCP Client, and MCP Server interfaces, and provide function calling functions.

10. The method according to claim 9, characterized by, The steps include intellectual property records: STA10: Insert the copyright vector space CV or interface of the intellectual property record in the entire VEM-sync vector sequence of the song file to store the copyright vector; STA20: The copyright vector includes copyright holder information, version number, authorization information, and encryption and decryption methods; STA30: Include encryption and decryption engines, including elliptic algorithms, public key, and private key modes; STA40: Managed using blockchain and cloud mode.

Citation Information

Patent Citations

  • VEM-Token beat capture and alignment model construction method

    CN120748450A

  • Speech and emotion synchronous recognition method based on neural network

    CN108806667A

  • Digital modulation and demodulation system and method based on hard synchronization

    CN117938611A

  • VEM-Token vocal music emotion multi-mode token song and accompaniment deep learning method

    CN120126506A

  • Automated methods and systems that infer, track, store, and report emotional states and emotional dynamics to render computational systems emotionally aware

    US12164680B1

Cited By

  • VEM-Token world model robot expression function construction method

    CN120951102A

  • Method for constructing robot expression function of vem-token world model

    CN120951102B

  • VEM-robot craniofacial end-to-end emotion bionic system

    CN121706840A