VEM-Token Vocal Emotion Multimodal Tokenization Deep Learning Method for Singing Voice and Accompaniment

Through the VEM-Token vocal emotions multimodal tokenization method, VEM classification and coordinate system are constructed, combined with supervised learning and deep learning, the singing stream and accompaniment stream are separated, and the multimodal emotion score and music score are generated, which solves the problem of not being able to identify vocal emotions in the existing technology, and realizes the multimodal input of vocal emotions and the development of agent Agent.

CN120126506BActive Publication Date: 2025-07-18GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510609148.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-18
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The prior art cannot effectively identify vocal emotions and perform multimodal inputs, and traditional NLP-Tokenization cannot adapt to multimodal inputs and analysis of vocal emotions.

Method used

The multimodal tokenization method of VEM-Token vocal emotions is used to construct VEM classification, VEM coordinate system and VEM function, combined with supervised learning and deep learning, vocal flow and accompaniment flow are separated, and multimodal emotion scores and music scores are generated.

Benefits of technology

Multimodal recognition and input of vocal emotions is realized, and the development of agent Agent is supported, and the problem of not being able to recognize vocal emotions in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126506B_ABST
    Figure CN120126506B_ABST
Patent Text Reader

Abstract

The VEM-Token vocal emotion multi-modal tokenization method for singing voice and accompaniment is a deep learning method that is different from the existing artificial intelligence methods which first segment information into literal token lemmas and then perform recognition. In the present invention, the vocal music file is spectrogrammed, the beat is detected, and the spectrogrammed vocal music file is segmented into a VEM-Token sequence according to the vocal music beat. A VEM coordinate system, VEM functions, and a VEM library are established according to multiple modalities such as lyrics, singing voice, accompaniment, singer emotion, accompaniment emotion, video, and image, and VEM-Token recognition is performed to separate the singing voice stream and the accompaniment stream. According to vocal music experts, a multi-modal emotion score is given to the vocal music sample, and VEM parameters are obtained by using supervised learning and deep learning algorithms to learn the multi-modal emotion of the vocal music sample. For other vocal music works, the multi-modal emotion of the vocal music can be recognized, and the lyrics score, VEM-Token singing voice score, VEM-Token accompaniment score, and VEM-Token music score can be output. It is connected to AI systems including common large models and developed into a vocal music intelligent agent Agent that can listen to music and recognize sheet music.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, specifically to the multi-modal recognition of vocal emotions. It adopts the deep learning method of VEM-Token for multi-modal tokenization of vocals and accompaniments in vocal emotions, solving the problems that current artificial intelligence cannot recognize vocal emotions and cannot effectively input music and singing, and providing an innovative solution for further realizing the intelligent agent Agent. Background Art

[0002] Token, in the early days in the field of network communication, referred to "token". In the field of artificial intelligence, it usually refers to the smallest unit or basic element in text processing. Currently, the identification methods of artificial intelligence for the real world mainly rely on Natural Language Processing (NLP) and Large Language Model (LLM), converting the language of humans into digital tokens that can be recognized by a computer in the form of text. For example, for Chinese or foreign characters and foreign letters, they are all converted into tokens, that is, the characters and letters are converted into binary codes, and this is used as the smallest identification unit. In fact, a token usually refers to a unit in text, which can be a word, a punctuation mark, a number, or even the input conversion of a sub-word.

[0003] It should be noted that "inputting" information into a computer and "outputting" information by the computer are completely different. Inputting is much more difficult than outputting because inputting requires the computer to understand. For example, for the input of "up, down, left, right", a strictly defined "dictionary" is needed to help the computer understand. As for "outputting", since the computer has already solved the understanding problem during "inputting", even if there is some processing in AI, the meaning during outputting is also well-founded.

[0004] Modality (in English: Modality) is the data form of the information type that an artificial intelligence system can process and understand, and the most basic one is text. And multi-modal (in English: Multimodal) refers to that an artificial intelligence model can simultaneously process and understand multiple types of data inputs, such as text, images, audio, video, etc., and achieve cross-modal interaction and generation.

[0005] So far, due to the relatively strict dictionary for defining and explaining text, the input of this modality based on text tokenization (that is, the conversion of text input into tokens) can still be accepted. However, for the modal forms of music and songs, there is no accurate tokenization input solution in the field of artificial intelligence (of course, for the output of music, that's another matter).

[0006] Furthermore, for a vocal song, it is important to accurately identify the emotion of the song, including the singer's voice and the sound of the accompanying instruments. In terms of emotion theory, based on the Ekman theory, the basic emotions of humans can be divided into: joy, sadness, anger, fear, calm, anticipation, love, hate, affection, and hatred. After expansion, it also includes related elements such as singing style, music style, singing method genre, and accompanying instruments. We classify all of these as "emotions". Obviously, this is also in the form of multi-modal. This invention application attempts to solve the learning method of multi-modal tokenization of vocal emotions and identify these multi-modal emotions into a new token.

[0007] To make the tokens distinguishable, this patent refers to the existing tokens as NLP-Token (Natural-Language-Processing Token) and simply refers to the multi-modal token of vocal emotion of this invention as VEM-Token (Vocal-Emotion-Multimodal Token).

[0008] Deficiencies of the existing technical methods

[0009] Based on the above analysis, the inventor believes that the existing text-based tokenization has the following deficiencies in vocal recognition and tokenization.

[0010] 1. It has not achieved multi-modal classification of vocal emotions and integration with AI.

[0011] 2. Traditional NLP-Tokenization cannot adapt to multi-modal input and parsing of vocal emotions.

[0012] 3. Traditional NLP-Tokenization cannot connect vocal emotions with NLP and LLM. Summary of the invention

[0013] According to the deficiencies of the existing technology, this invention proposes a brand-new and innovative method to realize the definition of VEM-Token and the deep learning method for multi-modal tokenization of vocal emotions, so as to solve the difficulty of conventional text-based NLP-Token methods for multi-modal vocal emotion input. Through the classification of multi-modal emotions, the establishment of a multi-modal emotion coordinate system and the design of a multi-modal emotion function, as well as through the evaluation of vocal samples by human vocal experts, supervised learning is established, and then the architecture and mechanism of deep learning are established to achieve the purpose and intention of this invention.

[0014] The purpose and intention of this invention are realized through the working steps of the following technical solutions.

[0015] 1. Implementation steps of the VEM basic solution

[0016] The present invention, as a VEM-Token vocal emotion multi-modal tokenized singing voice and accompaniment deep learning method, at least includes but is not limited to the following steps:

[0017] S100: Record emotions using one or more modalities, mark the vocal emotion multi-modal as VEM, and construct a VEM classification, a VEM coordinate system, a VEM function, and a VEM library.

[0018] S200: Collect vocal samples according to the VEM classification, have human vocal experts judge the emotions of the vocal samples in terms of singing voice and accompaniment, and use supervised learning and deep learning to train the VEM function to obtain VEM parameters and add them to the VEM library.

[0019] S300: Use a VEM processor to calibrate the beats of the vocal file, separate the singing voice stream and the accompaniment stream, perform VEM-Token segmentation on the vocal file according to the beats, convert the singing voice stream into a VEM-Token1 sequence, convert the accompaniment stream into a VEM-Token2 sequence, and add them to the preprocessing library.

[0020] S400: Use deep learning to generate a lyrics score, a VEM-Token singing voice score, a VEM-Token accompaniment score, and a VEM-Token music score respectively.

[0021] 2. Implementation steps of VEM emotion classification and VEM multi-modal emotion coordinate system

[0022] On the basis of the foregoing basic solution, in terms of the implementation steps of multi-modal emotion classification and multi-modal emotion coordinate system, the present invention specifically includes but is not limited to one or more combinations of the following steps or methods:

[0023] S110: Decompose emotions into VEM classifications based on one or more modalities and construct a VEM coordinate system, specifically including but not limited to:

[0024] S111: Independent emotions, including but not limited to the composition where emotions are independent of each other and have no association, construct a one-way one-dimensional coordinate axis, with the lowest point of the emotion as the coordinate 0 point and the highest point of the emotion as the maximum coordinate point.

[0025] S112: Opposite emotion pairs, including but not limited to the composition where two emotions are opposite to each other, construct a two-way one-dimensional coordinate axis, where the midpoint of the opposite emotion pair is the coordinate 0 point, the highest point of the positive emotion is the positive maximum coordinate point, and the highest point of the negative emotion is the negative maximum coordinate point.

[0026] S113: Associate opposite emotion groups, which are composed of opposite emotion pairs related between two or more than two opposite emotion pairs. Align their respective coordinate 0 points, make their respective two-way one-dimensional coordinate axes super-orthogonal, divide them with a hyperplane, place their respective positive emotions on the same side of the hyperplane, place their respective negative emotions on the opposite side of the hyperplane, construct a super-orthogonal coordinate system, take the highest point of their respective positive emotions as the positive maximum point of the super-orthogonal coordinates, and take the highest point of their respective negative emotions as the negative maximum point of the super-orthogonal coordinates.

[0027] The modality includes but is not limited to one or a combination of lyrics, singing voice, accompaniment, vocal style, music, emotional basis, accompanying instruments, video, and image.

[0028] The emotional basis includes but is not limited to one or a combination of joy, sadness, anger, fear, disgust, surprise, calmness, expectation, trust, love, hate, affection, and hatred.

[0029] The vocal style includes but is not limited to one or a combination of ethnic song singing methods, popular song singing methods, Western song singing methods, pop song singing methods, original song singing methods, and opera singing methods.

[0030] 3. Implementation steps of the VEM multi-modal emotion function

[0031] Based on the foregoing basic solution, in terms of the implementation steps of the multi-modal emotion function of the present invention, it specifically includes one or more combinations of the following steps or methods:

[0032] S120: Construct a VEM function, including but not limited to:

[0033] S121: For the one-way one-dimensional coordinate system of independent emotions, mark emotion scales on the one-way one-dimensional coordinate axis according to the attributes of the independent emotions, and construct an independent emotion VEM function.

[0034] S122: For the two-way one-dimensional coordinate system of opposite emotion pairs, mark emotion scales on the two-way one-dimensional coordinate axis according to the attributes of the opposite emotion pairs, and construct an opposite emotion pair VEM function.

[0035] S123: For the super-orthogonal coordinate system of associated opposite emotion groups, construct an associated opposite emotion group VEM function according to the attributes between the opposite emotion pairs included but not limited to within the associated opposite emotion groups, and the VEM projection relationship between the emotion scales on one two-way one-dimensional coordinate axis and the emotion scales on another two-way one-dimensional coordinate axis.

[0036] S124:

[0037] The VEM function attributes include but are not limited to the numerical values of the emotion scales of emotions on the coordinate axis.

[0038] The VEM functional relationships include, but are not limited to, the calculation methods between the numerical values of the emotion scales, and the calculation methods include, but are not limited to, one or a combination of the switching function, linear function, non-linear function, trigonometric function, and custom function between emotions.

[0039] The VEM projection relationships include, but are not limited to, one or a combination of trigonometric functions and custom functions.

[0040] S125: Add the VEM classification, VEM coordinate system, and VEM functions to the VEM library.

[0041] 4. Steps of Vocal Expert Evaluation and Supervised Learning

[0042] Based on the foregoing basic solution, in terms of the implementation steps of vocal expert evaluation and supervised learning of the present invention, it specifically includes, but is not limited to, one or a combination of the following steps or methods:

[0043] S210: Collect vocal samples. The vocal samples include, but are not limited to, recordings of songs with more than one VEM classification and more than one vocal style. Each vocal sample is evaluated by more than one human vocal expert, and emotional judgment scores are given respectively for the singing voice and the accompaniment. The scores and the vocal samples are added to the VEM library.

[0044] S220: The emotional judgment is obtained through steps including, but not limited to, S221, S222, and S223, specifically including, but not limited to.

[0045] S221: For independent emotions, based on the coordinate 0 point to the maximum coordinate point, the vocal expert gives scores respectively for the singing voice and the accompaniment.

[0046] S222: For opposite emotion pairs, based on the coordinate 0 point, the positive maximum coordinate point, and the negative maximum coordinate point, the vocal expert gives scores respectively for the singing voice and the accompaniment.

[0047] S223: For associated opposite emotion groups, based on the coordinate 0 point, the positive maximum hyper-orthogonal coordinate point, and the negative maximum hyper-orthogonal coordinate point of the hyper-orthogonal coordinate system, the vocal expert gives scores respectively for the singing voice and the accompaniment.

[0048] S230: Based on the vocal samples and scores in the VEM library, use supervised learning and deep learning to train the VEM functions and the VEM parameters included but not limited to by the VEM functions, and add the scores and VEM parameters to the VEM library.

[0049] 5. VEM Processor and VEM-Token Beat

[0050] Based on the foregoing basic solution, in terms of one of the hierarchical learning steps of the learning processor in the present invention, it specifically includes one or more combinations of the following steps or methods:

[0051] It includes the VEM processor to perform hierarchical processing on the vocal music file, and specifically further includes but is not limited to:

[0052] L1.0 layer processing: Use the VEM processor to convert vocal music files with more than one type and more than one channel into spectral format files in a unified format. Among them,

[0053] If the vocal music file is an analog signal file, it is converted into a spectral format file by analog-to-digital conversion using the sampling frequency.

[0054] If the vocal music file is a digital signal file, it is converted into a spectral format file, and the sampling frequency includes but is not limited to 44.1 kHz, 48 kHz, or an integer multiple of 44.1 kHz and 48 kHz.

[0055] L1.1 layer processing: If the spectral format file includes more than two channels, it is converted into a single-channel spectral format file of L1.1 layer.

[0056] L1.1.1 layer processing: Perform operations including but not limited to Fourier transform, constant-Q transform, Mel spectrogram transform, Hilbert transform, and discrete wavelet transform on the single-channel spectral format file of L1.1 layer to generate an L1.1.1 layer spectrogram.

[0057] L1.1.1.1 layer processing: For the L1.1.1 layer spectrogram, set up rhythm pointers to point to the detected spectral energy mutation points, or percussion rhythm points, or dynamic envelope rhythm points, or their combinations respectively, align the rhythm pointers, and pre-mark with the rhythm pointers as the beat starting point.

[0058] L1.1.2 layer processing: For the complete L1.1 layer, perform deep learning using a convolutional recurrent neural network, etc., to circularly detect the rhythm pointers. Those with pre-marked beat points coinciding are confirmed as beat starting point marks, and those not coinciding are marked as variable rhythm starting point marks. Starting from the variable rhythm starting point marks, execute the L1.1.1.1 layer processing in a loop to confirm and obtain the beat sequence of the complete vocal music file.

[0059] L1.2 layer processing: Based on the beat starting point mark, divide the spectral format file of the entire vocal music file into one VEM-Token according to the content of the spectral format file in each beat. According to the beat sequence, divide the entire spectral format file into a VEM-Token sequence and add it to the preprocessing library.

[0060] 6. VEM Processor and VEM-Token Musical Score

[0061] Based on the above basic solution, in terms of the second aspect of the hierarchical learning steps of the learning processor of the present invention, it specifically includes one or more combinations of the following steps or methods:

[0062] L1.4 layer processing: For L1.1 layer and L1.2 layer, set the beat intensity mark and beat type mark, take VEM-Token as the unit, detect the spectral energy, and divide it into pre-strong beat, pre-sub-strong beat, and pre-weak beat according to the spectral energy intensity.

[0063] L1.4.1 layer processing: For the complete L1.1 layer, adopt including but not limited to convolutional recurrent neural network and Bayesian model to conduct beat intensity detection and statistics, and circularly detect and mark pre-strong mark, pre-sub-strong beat, and pre-weak mark.

[0064] If the probability of the appearance of the pre-sub-strong beat is lower than the probability determination value, it is determined that the beat type mark of the vocal music file is including but not limited to 2 / 4 beat or 3 / 4 beat or 3 / 8 beat, that is, the strong, weak, and strong, weak, weak beat types.

[0065] If the probability of the appearance of the pre-sub-strong beat is higher than the probability determination value, it is determined that the beat type mark of the vocal music file is including but not limited to 4 / 4 beat, that is, the strong, weak, sub-strong, weak beat type, and the probability determination value is selected from 1% to 25%.

[0066] During the circular detection, if the pre-strong mark, pre-sub-strong beat, and pre-weak mark in the front and back sequences coincide, it is confirmed as the continuous rhythm type; if they do not coincide, it is marked as becoming strong and weak points, and start to circularly execute the L1.4 layer processing from the becoming strong and weak points to confirm the strong and weak sequence of the complete vocal music file.

[0067] L1.5 layer processing: For the vocal music file, mark the beat type and measure mark, and add them to the preprocessing library. The beat type includes at least one of 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, and 3 / 8 beat.

[0068] L1.6 layer processing: Based on L1.1 layer, L1.2 layer, and L1.5 layer, measure the frequency for VEM-Token, calculate and output the key signature according to the twelve-tone equal temperament in music theory and the rules of numbered musical notation and staff notation, and adopt including but not limited to convolutional recurrent neural network for deep learning to circularly detect the key signature. If a key change is found, record the new key signature and key change mark and add them to the preprocessing library.

[0069] L1.7 layer processing: According to the twelve-tone equal temperament rule in music theory and the spectral format file, calculate the pitch name sequence 1 existing in each VEM-Token, calculate the pitch name sequence 2 for the VEM-Token sequence, and construct the pitch name sequence 3 for the entire vocal music file and add it to the preprocessing library.

[0070] Among them, the pitch names are the symbols for recording pitches in the twelve-tone equal temperament in music theory, and the pitch name sequence is the musical score for recording pitch names and durations.

[0071] 7. VEM-Token Mono Processing

[0072] Based on the above basic solution, the present invention generates a multi-modal tokenized musical score for vocal emotions, specifically including one or more combinations of the following steps or methods:

[0073] Mono Processing Step:

[0074] S330: For a mono vocal music file, use a VEM processor to perform, according to the spectral format file, the preprocessing library, and the content of the VEM library, methods including but not limited to deep learning, specifically including but not limited to the encoding and decoding architectures of Dual-Stream U-Net, Demucs, or Spleeter in convolutional neural networks. Under the synchronization constraint of the beat sequence, perform multi-layer convolutional downsampling on the encoder and multi-layer convolutional upsampling on the decoder, decompose it into a vocal stream and an accompaniment stream, and add them to the preprocessing library.

[0075] S340: According to the beat sequence, decompose the vocal stream into a VEM-Token1 and a VEM-Token1 sequence, and decompose the accompaniment stream into a VEM-Token2 and a VEM-Token2 sequence, and add them to the preprocessing library.

[0076] S350: Synthesize the vocal stream and the accompaniment stream of the complete spectral format file into an L2.0 layer accompaniment stream file and an L3.0 layer vocal stream file, and add them to the preprocessing library.

[0077] 8. VEM-Token Multi-Channel Processing

[0078] Based on the above basic solution, the present invention also includes but not limited to the generation of high-frequency token0 segmentation, specifically including one or more combinations of the following steps or methods:

[0079] Multi-Channel Processing Step:

[0080] S360: For a multi-channel vocal music file, the steps include but are not limited to S361, S362, and S363, specifically:

[0081] S361: For vocal music files including but not limited to 2-channel, 4-channel, 5-channel and 5.1-channel, use the file content of each channel, convert it into a spectral format file, obtain the VEM-Token and VEM-Token sequence of each channel, and determine the VEM-Token1 and VEM-Token1 sequence, and the VEM-Token2 and VEM-Token2 sequence according to the format of the vocal music file.

[0082] S362: For vocal music files that are not 2-channel, 4-channel, 5-channel and 5.1-channel, use the file content of each channel, convert it into a spectral format file, obtain the VEM-Token and VEM-Token sequence of each channel, and determine the VEM-Token1 and VEM-Token1 sequence, and the VEM-Token2 and VEM-Token2 sequence according to the format of the vocal music file.

[0083] S363: Use a loss function for training, and split to obtain the L4.0 layer singing stream file and the L5.0 layer accompaniment stream file.

[0084] 9. VEM-Token Lyrics and Singing Score

[0085] Based on the foregoing basic solution, the present invention further includes one or more combinations of the following steps or methods:

[0086] S410: Use speech recognition to recognize the VEM-Token1 sequence, form lyrics, align the lyrics and the beat start marker, and split to obtain the lyric scores of the L1.8 layer single sentence lyrics, the L1.9 layer single paragraph lyrics, and the L1.A layer full song lyrics.

[0087] S420: According to the single sentence lyrics, single paragraph lyrics and full song lyrics, and according to the VEM library, use methods including but not limited to convolutional recurrent neural network and Bayesian model to cyclically detect VEM classification and calculate the distribution probability in the VEM-Token1 and VEM-Token1 sequences respectively, upgrade the VEM parameters, and add them to the preprocessing library.

[0088] S430: For a complete vocal music file, according to the VEM classification and distribution probability, align the lyrics and the beat start marker, calculate and output the VEM classification with the largest and the second largest probability in the distribution probability, and form the VEM-Token singing score.

[0089] 10. VEM-Token Instruments and Accompaniment Score

[0090] Based on the foregoing basic solution, the present invention further includes one or more combinations of the following steps or methods:

[0091] S440: Build an instrument sound library and add it to the VEM library.

[0092] S450: Based on the VEM library, using methods including but not limited to convolutional recurrent neural networks and Bayesian models, repeatedly detect the instrument sound library in VEM-Token2 and VEM-Token2 sequences, calculate the matching probability, and add it to the preprocessing library.

[0093] S460: For a complete vocal music file, based on the VEM classification and instrument matching probability, align the lyrics and the beat start markers, calculate and output the VEM classification with the highest probability among the matching probabilities, and calculate and obtain the VEM-Token accompaniment scores of the numbered musical notation and the staff notation according to the pitch name sequence 3 and the rules of music theory.

[0094] S470: Based on the VEM-Token accompaniment score, aligning with the beat start marker, merge the VEM-Token vocal score, the VEM-Token accompaniment score, and the lyrics score into a VEM-Token music score.

[0095] 11. Connect to NLP and LLM

[0096] Based on the above solutions, the present invention further includes one or more combinations of the following steps or methods:

[0097] Access the resources of natural language processing NLP or large language model LLM according to the lyrics or the lyrics score to generate a vocal emotion score.

[0098] Access the resources of natural language processing NLP or large language model LLM according to the accompaniment stream and the VEM-Token accompaniment score to generate an accompaniment emotion score.

[0099] Synthesize the vocal emotion score and the accompaniment emotion score to obtain a synthesized music score.

[0100] According to the communication protocol, connect the VEM-Token, the VEM library, and the preprocessing library to a common large model system to customize a vocal music intelligent agent Agent.

[0101] 12. Invention purpose and intention

[0102] Through long-term research, observation and experiments, the inventor proposed a deep learning method for VEM-Token vocal emotion multi-modal tokenization, which is a new method for artificial intelligence to recognize and input vocal emotion multi-modalities.

[0103] The purpose and intention of the present invention are:

[0104] 1. Innovatively establish a multi-modal classification of VEM-Token vocal emotions.

[0105] Including but not limited to multi-modal emotion classification, multi-modal emotion coordinates, and multi-modal emotion functions, which are different from traditional NLP-based techniques that are separated by text tokens. They are underlying tokenization techniques that are perfectly suitable for the brand-new vocal emotion multi-modal.

[0106] 2. Innovate a new multi-modal tokenization method for VEM-Token vocal emotion.

[0107] The inventor believes that the smallest unit at the bottom layer of vocals should be based on the segmentation of musical beats, rather than text-based segmentation. Because vocals are different from articles. Articles can use words as the smallest unit to express the content of the article, while vocals use sounds as the smallest unit to express the meaning of the song.

[0108] 3. Establish a supervised learning library.

[0109] Based on the multi-modal classification of vocal emotions, collect vocal samples, have human vocal experts judge the vocal samples, establish a sample library, perform supervised learning, further guide the artificial intelligence to identify vocals, and finally establish a multi-modal emotion library.

[0110] 4. VEM-Token supports access to existing NLP and LLM resources to implement an intelligent agent, Agent.

[0111] Adopt the multi-modal tokenization method of vocal emotion, perform in-depth learning by AI to complete the emotion recognition of vocal files, and then access NLP and LLM so that the existing NLP and LLM resources can support the development of the intelligent agent, Agent, for multi-modal vocal emotions.

[0112] 13. Beneficial effects of the invention

[0113] 1. The invention purpose and intention are achieved.

[0114] 2. Solve the problem that existing artificial intelligence cannot recognize multi-modal vocal emotions.

[0115] 3. Support the implementation of the music intelligent agent, Agent. Brief description of the drawings

[0116] List of drawings:

[0117] Figure 1 : System flow schematic diagram

[0118] Figure 2 : Multi-modal independent emotion and coordinate diagram

[0119] Figure 3 : Opposite emotion pair and coordinate diagram

[0120] Figure 4: Associated Opposite Emotion Group and Hyper-orthogonal Coordinate Diagram

[0121] Figure 5 : Schematic Diagram of VEM-Token Segmentation

[0122] Figure 6 : VEM-Token Musical Score Diagram

[0123] Detailed Description of the Drawings:

[0124] See the specific embodiments for details. Specific Embodiments

[0125] The objectives and intentions of the present invention can be achieved by the following specific embodiments. It should be particularly noted here that since the specific embodiments all have specific uses and industrial applicability, the embodiments cannot include all the features and steps of the present invention, nor are they a limitation to the present invention. The description in the claims of the present invention is the summary of the invention.

[0126] This example is a general example of the present invention. It should be stated that: the content and views of this embodiment are not a limitation to the present invention, nor can they fully interpret the present invention. It is only one of the embodiments of the application of the present invention in the field of AI.

[0127] The specific embodiments of the present invention are as follows:

[0128] An AI system that can "listen to music" and "read sheet music"

[0129] ——An innovative VEM-Token vocal emotion multi-modal tokenization method for vocal and accompaniment deep learning.

[0130] Illustration Explanation

[0131] The content of this embodiment mainly includes but is not limited to the following main schematic drawings, which are: Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 、 Figure 5 、 Figure 6 。

[0132] Implementation Step Explanation

[0133] The method steps of this embodiment mainly include Step 1 to Step 11. Among them, unless otherwise specified, the step numbers of these 11 parts do not have a sequential order, nor do all embodiments require the combination of these 11 parts. In addition, each of these 11 parts includes several sub-steps. Unless otherwise specified, these sub-steps are not completely required, and their sequential order is not necessary. Instead, they are optimized and further selected by the patent implementer according to the requirements of some specific tasks.

[0134] The specific working steps are described as follows:

[0135] 1. Implementation steps of the VEM basic solution

[0136] As a VEM-Token vocal emotion multimodal tokenization method for vocal music and accompaniment deep learning, the present invention includes but is not limited to the following steps:

[0137] S100: Record emotions using one or more modalities, mark the vocal emotion multimodal as VEM, and construct a VEM classification, a VEM coordinate system, a VEM function, and a VEM library.

[0138] S200: Collect vocal music samples according to the VEM classification, have human vocal music experts judge the emotions of the vocal music samples in terms of singing and accompaniment, and use supervised learning and deep learning to train the VEM function to obtain VEM parameters and add them to the VEM library.

[0139] S300: Use a VEM processor to calibrate the beats of the vocal music file, separate the singing stream and the accompaniment stream, perform VEM-Token segmentation on the vocal music file according to the beats, convert the singing stream into a VEM-Token1 sequence, convert the accompaniment stream into a VEM-Token2 sequence, and add them to the preprocessing library.

[0140] S400: Use deep learning to generate a lyrics score, a VEM-Token singing score, a VEM-Token accompaniment score, and a VEM-Token music score respectively.

[0141] Figure 1 It is a schematic diagram of the system process of the present invention.

[0142] In Figure 1 , starting from the left, there are the input of the vocal music file, multimodal emotion classification, vocal music samples, and vocal music experts respectively. On the far right is

[0143] It should be particularly noted that Figure 1 The content, position, and connection in the boxes in

[0144] It should be particularly noted that the "tokenization" in the present invention refers to, in accordance with the common terminology in the current field of artificial intelligence, that a token refers to the "lexicalization" of human language texts based on NLP. That is to say, for language texts, with the smallest unit of characters, such as letters and words in the Western language system and Chinese characters in Chinese, artificial intelligence divides the article into individual "lexical tokens" and inputs them into the AI system, and then combines them to form sentences, paragraphs, and articles. Therefore, "tokenization" is the process of decomposing an article into the smallest units.

[0145] So far, tokenization has been able to handle the text information modality well. However, for other information modalities, such as spoken language and speech, they are still converted into text for input. For images and videos, in fact, tokenization and input are still completed through web pages or descriptive texts. Although for the output of artificial intelligence, such as Microsoft's Muse AI, it can output and create music and vocals well, unfortunately, it still cannot complete the input of vocals. All of this is because so far, predecessors have not found a way for tokenized input of vocals.

[0146] Figure 1 It includes the following general steps:

[0147] 1. Vocal emotion multi-modal classification (VEM classification) is an initial VEM classification that needs to be pre-constructed in the present invention. Since the present invention adopts the dynamic update mechanism of artificial intelligence, as the system runs, the VEM classification will be gradually iterated and updated. Through the initial classification, VEM coordinates and VEM functions are established, and these rules are input and archived into the VEM library.

[0148] 2. Vocal samples are some song samples that are manually collected based on the initial VEM classification and are typically representative in terms of classification. These samples are judged by human vocal experts based on their own experience, and scores for the VEM singing parameters and VEM accompaniment parameters are given. It should be noted that the setting of these scores has a pre-set range, such as 0 to 10 points, and there are also pre-set scoring rules for these scores. For example, what each segment means. On the software interface, according to the pre-set segment scoring instructions, subjective scoring is performed by vocal experts. These scores are converted into VEM parameters in the VEM function and incorporated into the VEM library through supervised learning set by the system.

[0149] 3. The vocal music files are the songs fed into the system, including files in relevant formats, such as but not limited to MP3, MPEG, MPEG-4, WAV, WMA, FLAC, AAC, OGG, MIDI and other formats. These vocal music files in these formats are converted by the system into spectrum formatted files with a unified format for the system to identify. Then, after hierarchical processing, they are sent to the VEM processor for the next step of processing.

[0150] 4. The role of the VEM library is also a dynamic emotion database storing the relevant rules and parameters of the system. When the subsequent learning processor runs, in the spectrum formatted files after the conversion of the vocal music files, the music beats are divided, and according to the music beats, they are segmented into the VEM-Token of the present invention, completed tokenization, and sent to the preprocessing library for storage. After deep learning, it is decomposed into the vocal music stream VEM-Token1 and the accompaniment stream VEM-Token2, and finally forms the VEM-Tokenized musical score.

[0151] In fact, since the most core innovation of the present invention is to tokenize the vocal music files based on multi-modal vocal music emotions, that is, VEM-Token, to solve the problem that the NLP-token based on natural language processing in the existing artificial intelligence field cannot complete the input and recognition of multi-modal vocal music emotions. The proposal of the VEM-Token of the present invention solves the tokenization of input and recognition from the bottom layer. In addition, after the VEM-Token is connected to artificial intelligence, the existing AI technology can well realize the understanding of vocal music, and then various subsequent problems can be solved by means of the existing artificial intelligence technology. For example, more and better analysis and musical score output.

[0152] 2. Implementation steps of VEM emotion classification and VEM multi-modal emotion coordinate system

[0153] On the basis of the foregoing basic solution, in terms of the implementation steps of multi-modal emotion classification and multi-modal emotion coordinate system, the present invention specifically includes one or more combinations of the following steps or methods:

[0154] S110: Based on more than one modality, decompose emotions into VEM classifications and construct a VEM coordinate system, specifically including but not limited to:

[0155] S111: For independent emotions, including but not limited to the composition where emotions are independent of each other and have no association, construct a one-way one-dimensional coordinate axis, with the lowest point of the emotion as the coordinate 0 point and the highest point of the emotion as the maximum coordinate point.

[0156] Preferably, S112: Opposite emotion pairs, including but not limited to two components with opposite emotions to each other, construct a two-way one-dimensional coordinate axis. Among them, the midpoint of the opposite emotion pair is the coordinate 0 point, the highest point of the positive emotion is the positive maximum point of the coordinate, and the highest point of the negative emotion is the negative maximum point of the coordinate.

[0157] Preferably, S113: Associated opposite emotion groups, including opposite emotion pairs related between more than one group of two opposite emotion pairs, are aligned with their respective coordinate 0 points coinciding, their respective two-way one-dimensional coordinate axes are hyper-orthogonal, divided by a hyperplane boundary, with their respective positive emotions placed on the same side of the hyperplane and their respective negative emotions placed on the opposite side of the hyperplane, to construct a hyper-orthogonal coordinate system. The highest point of their respective positive emotions is the positive maximum point of the hyper-orthogonal coordinate, and the highest point of their respective negative emotions is the negative maximum point of the hyper-orthogonal coordinate.

[0158] The modality includes but is not limited to one of lyrics, singing voice, accompaniment, vocal style, music, emotional basis, accompanying instruments, video, and image. Multimodality is a combination of these modalities.

[0159] The emotional basis includes but is not limited to one or a combination of joy, sadness, anger, fear, disgust, surprise, calmness, expectation, trust, love, hate, affection, and hatred.

[0160] The vocal style includes but is not limited to one or a combination of ethnic song singing methods, pop song singing methods, Western song singing methods, popular song singing methods, original song singing methods, and opera singing methods.

[0161] It should be noted that independent emotions, opposite emotion pairs, and associated opposite emotion groups are sometimes relative. For vocal works of different styles and vocal works of different cultural backgrounds, the classification of modalities is not always the same. Therefore, in extreme cases, independent emotions, opposite emotion pairs, and associated opposite emotion groups all refer to a single vocal work.

[0162] It should be emphasized that the "hyper-orthogonal coordinate system" here refers to the arrangement of right-angled coordinate axes with more than 3 dimensions in the same system, such as 4D, 5D or more. Since we cannot directly draw it on paper, we use the hyper-dimensional space in mathematics to describe and record. In addition, the mutual correlation of these hyper-dimensions is not uniquely orthogonal. It may also be just intersecting in mathematics rather than orthogonal, or even be curved coordinates based on Riemannian geometry.

[0163] Figure 2 It is a multimodal independent emotion and coordinate diagram.

[0164] For example, emotions such as video, lyrics, vocal style, classical guitar, and accompaniment are independent of each other without any association. The inventor defines them as one-dimensional coordinate axes with their respective coordinate zero points coinciding. The highest points of each emotion are set as the maximum points of the coordinate axes, thus forming a multi-modal independent emotion coordinate system. It should be emphasized here that although the coordinate axes show angles, in fact, there is no association between the coordinate axes, and they are independent of each other. They may not be in the same plane, but are simply drawn together when making the graph. For example, there is no connection between lyrics, vocal style, and accompanying instruments, and one cannot be interpreted through another.

[0165] Figure 3 It is a graph of opposite emotion pairs and coordinates.

[0166] Some emotions are actually opposite, such as happiness and sadness, approval and suspicion, love and hate, joy and sorrow, affection and enmity. For such emotions, the inventor calls them opposite emotion pairs and places them at both ends of a one-dimensional coordinate axis respectively, with opposite directions, symmetric positive and negative scores, and the midpoint as the coordinate zero point. The highest points of each opposite emotion pair are set as the positive and negative maximum points of the coordinate axis. Opposite emotion pairs can be either independent of each other or related to each other.

[0167] Figure 4 It is a graph of related opposite emotion groups and hyper-orthogonal coordinates.

[0168] Among some opposite emotion pairs, there are associations, such as love and hate, joy and sorrow, affection and enmity. These opposite emotion pairs form related opposite emotion groups. Among them, a combination of emotions, such as Figure 4 point E in it, can be associated with the coordinate axes of love and hate, joy and sorrow, affection and enmity respectively. For example, point E is in a hyper-orthogonal coordinate system, and its respective coordinates for the love X-axis, sorrow -Y-axis, and affection Z-axis are E(x, y, z). Assuming a linear relationship between these 3 coordinate axes, their projections are as shown in the pink and light blue boxes in Figure 4 Accordingly, the inventor places them in a hyper-orthogonal coordinate system and can establish a mathematical model of mutual association. The reason it is called a hyper-orthogonal coordinate system is that there are more than 3 related opposite emotion pairs, forming a high-dimensional hyper-space.

[0169] 3. Implementation steps of the VEM multi-modal emotion function

[0170] Based on the above-mentioned solutions, in terms of the implementation steps of the multi-modal emotion function of the present invention, it specifically includes one or more combinations of the following steps or methods:

[0171] S120: Construct a VEM function, including but not limited to:

[0172] S121: For the one - dimensional coordinate system of independent emotions, mark the emotion scales on the one - dimensional coordinate axis according to the attributes of independent emotions, and construct the independent emotion VEM function.

[0173] Preferably, S122: For the two - dimensional one - dimensional coordinate system of opposite emotion pairs, mark the emotion scales on the two - dimensional one - dimensional coordinate axis according to the attributes of opposite emotion pairs, and construct the opposite emotion pair VEM function.

[0174] Preferably, S123: For the super - orthogonal coordinate system of associated opposite emotion groups, construct the associated opposite emotion group VEM function according to the attributes between the opposite emotion pairs included but not limited to within the associated opposite emotion groups, and the VEM projection relationship between the emotion scales on one two - dimensional one - dimensional coordinate axis and the emotion scales on another two - dimensional one - dimensional coordinate axis.

[0175] Preferably, S124 includes:

[0176] The VEM function attributes include but are not limited to the numerical values of the emotion scales of emotions on the coordinate axis.

[0177] The VEM function relationships include but are not limited to the calculation methods between the numerical values of the emotion scales. The calculation methods include but are not limited to one or a combination of the switching function, linear function, non - linear function, trigonometric function, and custom function between emotions.

[0178] The VEM projection relationships include but are not limited to one or a combination of trigonometric functions and custom functions.

[0179] S125: Add the VEM classification, VEM coordinate system, and VEM function to the VEM library.

[0180] The VEM function relationships are sometimes also fuzzy. We can also borrow the function relationships of fuzzy mathematics for simulation and analysis. For example, introduce the concept of fuzzy membership degree and use existing fuzzy functions to describe.

[0181] The VEM projection relationship such as Figure 4 Point E in it, which is composed of three emotions of "love, affection, worry". We can understand these three coordinate systems as linear scales, and further project point E onto the three coordinate axes of "love, affection, worry" in detail, generating three projections of (x, y, z). Obtain the emotion scale of the "affection" coordinate as x, the emotion scale of the "love" coordinate as y, and the emotion scale of the "worry" coordinate as z. Assuming that these three emotions conform to trigonometric function relationships, then through three - dimensional analytic geometry operations, a more detailed and accurate description of the numerical emotion scale can be obtained.

[0182] 4. Steps of vocal music expert evaluation and supervised learning

[0183] Based on the foregoing solution, in terms of the implementation steps of vocal music expert evaluation and supervised learning, the present invention specifically includes one or more combinations of the following steps or methods:

[0184] S210: Collect vocal music samples. The vocal music samples include, but are not limited to, recordings of songs with more than one VEM classification and more than one vocal music style. Each vocal music sample is evaluated by more than one human vocal music expert, and scores for emotional judgment are given respectively for the singing voice and the accompaniment. The scores and the vocal music samples are added to the VEM library.

[0185] Further, S220: The emotional judgment is obtained through steps including, but not limited to, S221, S222, and S223, specifically including, but not limited to.

[0186] S221: For independent emotions, based on the coordinate 0 point to the maximum coordinate point, vocal music experts give scores respectively for the singing voice and the accompaniment.

[0187] Preferably, S222: For opposite emotion pairs, based on the coordinate 0 point, the positive maximum coordinate point, and the negative maximum coordinate point, vocal music experts give scores respectively for the singing voice and the accompaniment.

[0188] Preferably, S223: For associated opposite emotion groups, based on the coordinate 0 point, the positive maximum hyper-orthogonal coordinate point, and the negative maximum hyper-orthogonal coordinate point of the hyper-orthogonal coordinate system, vocal music experts give scores respectively for the singing voice and the accompaniment.

[0189] Preferably, S230: Based on the vocal music samples and scores in the VEM library, supervised learning and deep learning are used to train the VEM function and the VEM parameters included but not limited to by the VEM function. The scores and the VEM parameters are added to the VEM library.

[0190] The steps here are actually to let people who understand vocal music mark the vocal music samples manually. It should be noted that since this marking work requires vocal music professional knowledge, not only a large number of typical vocal music samples need to be found, but also a wide coverage is required. At the same time, multiple professionals are needed to operate. In this way, the supervised learning work can be better completed to achieve the best learning effect and lay the best foundation for subsequent deep learning.

[0191] 5. VEM Processor and VEM-Token Beat

[0192] Based on the foregoing solution, in terms of one of the hierarchical learning steps of the learning processor, the present invention specifically includes one or more combinations of the following steps or methods:

[0193] It includes the VEM processor to perform hierarchical processing on the vocal music file, and specifically further includes, but not limited to:

[0194] Processing of Layer L1.0: Using a VEM processor to convert vocal music files with more than one type and more than one channel into a spectral format file in a unified format. Among them,

[0195] Preferably, if the vocal music file is an analog signal file, it is converted into a spectral format file after being converted from analog to digital using a sampling frequency.

[0196] Preferably, if the vocal music file is a digital signal file, it is converted into a spectral format file.

[0197] Furthermore, the sampling frequency includes but is not limited to 44.1 kHz, 48 kHz, or integer multiples of 44.1 kHz and 48 kHz.

[0198] Processing of Layer L1.1: If the spectral format file includes but is not limited to more than two channels, it is converted into a single-channel spectral format file of Layer L1.1.

[0199] Processing of Layer L1.1.1: Performing operations including but not limited to Fourier transform, constant Q transform, Mel spectrogram transform, Hilbert transform, and discrete wavelet transform on the single-channel spectral format file of Layer L1.1 to generate a spectral diagram of Layer L1.1.1.

[0200] Processing of Layer L1.1.1.1: For the spectral diagram of Layer L1.1.1, set up rhythm pointers to respectively point to the detected spectral energy mutation points, or percussion rhythm points, or dynamic envelope rhythm points, or combinations thereof.

[0201] Furthermore, align the rhythm pointers and pre-mark with the rhythm pointers as the beat starting points.

[0202] Processing of Layer L1.1.2: For the complete Layer L1.1, perform deep learning using, including but not limited to, a convolutional recurrent neural network to cyclically detect the rhythm pointers. Those with pre-marked beat points coinciding are confirmed as beat starting point marks, and those not coinciding are marked as variable rhythm starting point marks. Starting from the variable rhythm starting point marks, execute the processing of Layer L1.1.1.1 in a loop to confirm and obtain the beat sequence of the complete vocal music file.

[0203] Preferably, processing of Layer L1.2: According to the beat starting point marks, divide the spectral format file of the entire vocal music file into one VEM-Token according to the content of the spectral format file in each beat. According to the beat sequence, divide the entire spectral format file into a sequence of VEM-Tokens and add them to the preprocessing library.

[0204] Figure 5 It is a schematic diagram of VEM-Token segmentation.

[0205] In the figure, VEM-Token is a sequence of spectrum format files converted from vocal music files. Beat is the starting mark of the beat. VEM-Token1 is segmented from the vocal stream, and VEM-Token2 is segmented from the accompaniment stream. Among them, n, n+1, n+2, and n+3 are the serial numbers of the tokenized series respectively.

[0206] It should be emphasized that VEM-Token is actually composed of two parts. One is the division of VEM-Token, which is completed from the starting point of the beat to the next starting point of the beat. The other is the content of VEM-Token, such as Figure 5 As shown, VEM-Token, VEM-Token1, and VEM-Token2 are respectively composed of segments in the vocal music file sequence, vocal stream sequence, and accompaniment stream sequence. Therefore, the data dimensions of VEM-Token, VEM-Token1, and VEM-Token2 at least include beat marks and content.

[0207] In addition, the VEM processor is usually composed of software or process steps, and can also be composed of embedded hardware, including CPU and GPU.

[0208] 6. VEM Processor and VEM-Token Musical Score

[0209] Based on the foregoing solution, in the second aspect of the hierarchical learning steps of the learning processor of the present invention, it specifically includes one or more combinations of the following steps or methods:

[0210] L1.4 layer processing: For L1.1 layer and L1.2 layer, set the beat intensity mark and beat type mark. Taking VEM-Token as the unit, detect the spectral energy, and divide it into pre-strong beat, pre-sub-strong beat, and pre-weak beat according to the spectral energy intensity.

[0211] Further, L1.4.1 layer processing: For the complete L1.1 layer, adopt methods including but not limited to convolutional recurrent neural network and Bayesian model to perform beat intensity detection and statistics, and circularly detect and mark the pre-strong mark, pre-sub-strong beat, and pre-weak mark.

[0212] Preferably, if the probability of the pre-sub-strong beat appearing is lower than the probability determination value, it is determined that the beat type mark of the vocal music file is including but not limited to 2 / 4 beat or 3 / 4 beat or 3 / 8 beat, that is, strong, weak, and strong, weak, weak beat types.

[0213] Preferably, if the probability of the pre-sub-strong beat appearing is higher than the probability determination value, it is determined that the beat type mark of the vocal music file is including but not limited to 4 / 4 beat, that is, strong, weak, sub-strong, weak beat type. The probability determination value is selected from 1% to 25%.

[0214] Furthermore, during the loop detection, if the pre-strong mark, pre-secondary strong beat and pre-weak mark in the previous and subsequent sequences overlap, they are confirmed as a continuous rhythm type. If they do not overlap, they are marked as variable strong and weak points. The L1.4 layer processing is executed in a loop starting with the variable strong and weak points to confirm the strong and weak sequence of the complete vocal file.

[0215] Furthermore, L1.5 layer processing: for vocal files, the beat type and measure mark are marked and added to the preprocessing library, and the beat type includes at least but is not limited to one of 2 / 4, 3 / 4, 4 / 4, and 3 / 8.

[0216] Furthermore, the L1.6 layer processes: based on the L1.1 layer, the L1.2 layer and the L1.5 layer, the frequency of the VEM-Token is measured, and the output key signature is calculated according to the twelve-tone equal temperament in music theory and the rules of the simplified notation and the five-line notation. Deep learning is performed using methods including but not limited to convolutional recurrent neural networks to cyclically detect key signatures. If a key change is found, the new key signature and key change mark are recorded and added to the preprocessing library.

[0217] Furthermore, L1.7 layer processing: according to the twelve-tone equal temperament rule in music theory and the spectrum format file, for each VEM-Token, the note name sequence 1 existing therein is calculated, for the VEM-Token sequence, the note name sequence 2 is calculated, and for the entire vocal file, the note name sequence 3 is constructed and added to the preprocessing library.

[0218] The note names are symbols used in music theory to record the pitches of the twelve-tone equal temperament, and the note name sequence is a musical score that records the note names and duration.

[0219] It should be noted that, for the identification of 4 / 4 beats, judging by "secondary strong beat" is a conventional method according to music theory. However, in a vocal work, due to the disturbance of human singing, the basis for judgment is often unclear, so the present invention introduces probability statistics for further identification.

[0220] In addition, readers of this patent need to understand that it is important to have relatively professional vocal knowledge and music theory knowledge.

[0221] 7. VEM-Token mono processing

[0222] On the basis of the above scheme, the present invention generates multimodal tokenized music scores of vocal emotions, including but not limited to the following one or more combined steps or methods:

[0223] Mono processing steps:

[0224] S330: For monophonic vocal files, a VEM processor is used according to the spectrum format file, the preprocessing library and the VEM library content, using methods including but not limited to deep learning, specifically including but not limited to the dual-stream U-Nnt, Demucs or Spleeter encoding and decoding architecture in the convolutional neural network. Under the synchronization constraint of the beat sequence, the encoder is subjected to multi-layer convolution down-sampling, and the decoder is subjected to multi-layer convolution up-sampling, and the files are decomposed into a singing stream and an accompaniment stream, which are then added to the preprocessing library.

[0225] S340: According to the beat sequence, decompose the vocal stream into VEM-Token1 and VEM-Token1 sequence, decompose the accompaniment stream into VEM-Token2 and VEM-Token2 sequence, and add them to the preprocessing library.

[0226] Preferably, S350: synthesize the vocal stream and accompaniment stream of the complete spectrum format file into an L2.0 layer accompaniment stream file and an L3.0 layer vocal stream file, and add them to the preprocessing library.

[0227] Figure 6 This is the VEM-Token music score chart.

[0228] This is a simple schematic diagram of the operating results of the present invention, including but not limited to simple musical notation, five-line musical notation, lyrics, multimodal representation of emotions, etc. Among them, multimodal representation of emotions includes but is not limited to description of singing style, description of singing emotions, description of accompaniment instrument emotions and foil, etc. Users of this patent should understand that, inspired by the VEM-Token music score diagram of this invention, users of this patent can follow this idea and draw music scores with richer emotions and content without creating new work.

[0229] 8. VEM-Token multi-channel processing

[0230] On the basis of the above scheme, the present invention also includes but is not limited to high-frequency token0 segmentation generation, specifically including but not limited to one or more of the following combined steps or methods:

[0231] Multi-channel processing steps:

[0232] S360: For a multi-channel vocal music file, the steps include but are not limited to S361, S362 and S363, specifically:

[0233] S361: For vocal music files including but not limited to 2-channel, 4-channel, 5-channel and 5.1-channel, use the file content of each channel, convert it into a spectral format file, obtain the VEM-Token and VEM-Token sequence of each channel, and determine the VEM-Token1 and VEM-Token1 sequence, and the VEM-Token2 and VEM-Token2 sequence according to the format of the vocal music file.

[0234] Preferably, S362: For vocal music files that are not 2-channel, 4-channel, 5-channel and 5.1-channel, use the file content of each channel, convert it into a spectral format file, obtain the VEM-Token and VEM-Token sequence of each channel, and determine the VEM-Token1 and VEM-Token1 sequence, and the VEM-Token2 and VEM-Token2 sequence according to the format of the vocal music file.

[0235] Preferably, S363: Use a loss function for training to segment and obtain the L4.0 layer vocal music stream file and the L5.0 layer accompaniment stream file.

[0236] It should be emphasized that in the formats of many vocal music files, they already contain multiple channels, such as MP3, WAV, WMA, FLAC, AAC, OGG, etc. Some of them contain a single channel, and some contain multiple channels. When developing, the user of this patent should develop a software for vocal music file format recognition and channel recognition, run this software first, and then proceed with the next step.

[0237] 9. VEM-Token Lyrics and Vocal Music Score

[0238] Based on the foregoing solutions, the present invention further includes one or more combinations of the following steps or methods:

[0239] S410: Use speech recognition to recognize the VEM-Token1 sequence, form lyrics, align the lyrics and the beat start marker, and segment and obtain the lyric scores of the L1.8 layer single sentence lyrics, the L1.9 layer single paragraph lyrics, and the L1.A layer full song lyrics.

[0240] Preferably, S420: According to the single sentence lyrics, single paragraph lyrics and full song lyrics, and based on the VEM library, use including but not limited to convolutional recurrent neural network and Bayesian model to cyclically detect the VEM classification and calculate the distribution probability in the VEM-Token1 and VEM-Token1 sequence respectively, upgrade the VEM parameters, and add them to the preprocessing library.

[0241] Preferably, S430: For a complete vocal music file, align the lyrics and the beat start markers according to the VEM classification and the distribution probability, calculate and output the VEM classifications with the largest and the second largest probabilities in the distribution probability, and form a VEM-Token vocal music score.

[0242] It should be emphasized that the lyrics include different languages, such as Chinese, English, Italian, French, etc., and can also include various dialects, and even pinyin.

[0243] In addition, in terms of emotion judgment, similar to the repeated calculation of a recurrent neural network and the statistical evaluation of a probability model, it is beneficial to more accurately judge the emotion.

[0244] 10. VEM-Token Instrument and Accompaniment Score

[0245] Based on the foregoing solutions, the present invention further includes one or more combinations of the following steps or methods:

[0246] S440: Construct an instrument sound effect library and add it to the VEM library.

[0247] S450: According to the VEM library, using, including but not limited to, a convolutional recurrent neural network and a Bayesian model, respectively loop-detect the instrument sound effect library in the VEM-Token2 and VEM-Token2 sequences, calculate the matching probability, and add it to the preprocessing library.

[0248] Further, S460: For a complete vocal music file, align the lyrics and the beat start markers according to the VEM classification and the instrument matching probability, calculate and output the VEM classification with the largest probability in the matching probability, and calculate and obtain the VEM-Token accompaniment scores of the numbered musical notation and the staff notation according to the pitch name sequence 3 and the rules of music theory.

[0249] Further, S470: Align according to the beat start marker based on the VEM-Token accompaniment score, and merge the VEM-Token vocal music score, the VEM-Token accompaniment score, and the lyrics score into a VEM-Token music score.

[0250] In the figure, from bottom to top are the lyrics, the numbered musical notation, the staff notation, the VEM classification, and its emotion level. It should be noted that Figure 6 This is only a schematic diagram of a music score. Users can customize the personalized layout and software according to their own preferences.

[0251] 11. Connection with NLP and LLM

[0252] Based on the foregoing solutions, the present invention further includes one or more combinations of the following steps or methods:

[0253] According to the lyrics or the lyric score, access the resources of natural language processing NLP or large language model LLM to generate a singing emotion spectrum.

[0254] According to the accompaniment stream and the VEM-Token accompaniment score, access the resources of natural language processing NLP or large language model LLM to generate an accompaniment emotion spectrum.

[0255] Synthesize the singing emotion spectrum and the accompaniment emotion spectrum to obtain a synthesized music score.

[0256] According to the communication protocol, connect the VEM-Token, VEM library, and preprocessing library to the common large model system to customize the vocal music intelligent agent Agent.

[0257] It should be noted that under the innovation of VEM-Token proposed in the present invention, based on the principle of this "lexical token" division, connecting VEM-Token to the resources of natural language processing NLP or large language model LLM is conducive to obtaining the huge advantages of current artificial intelligence resources. For example, connecting to a similar common large model system to obtain a secondary development system. Further, according to the concept of intelligent agent Agent, the present invention supports the development into a vocal music intelligent agent.

Claims

1. VEM-Token Vocal Emotion Multimodal Tokenization Deep Learning Method for Singing Voice and Accompaniment, characterized in that, Including: S100: Record emotions using more than one modality, mark the vocal emotion multimodality as VEM, and construct a VEM classification, a VEM coordinate system, a VEM function, and a VEM library; the emotions include one or a combination of joy, sadness, anger, fear, disgust, surprise, calm, anticipation, trust, love, hate, affection, and hatred, and the modalities include one or a combination of lyrics, singing voice, accompaniment, vocal style, music, emotional basis, accompanying instruments, video, and image. The VEM coordinate system includes an axis system established based on independent emotions, opposite emotion pairs, and associated opposite emotion groups. S200: Collect vocal samples according to the VEM classification, have human vocal experts judge the emotions of the vocal samples in terms of singing voice and accompaniment, and use supervised learning and deep learning to train the VEM function to obtain VEM parameters and add them to the VEM library. S300: Use a VEM processor to calibrate the beats of a vocal file, separate the singing voice stream and the accompaniment stream, perform VEM-Token segmentation on the vocal file according to the beats, convert the singing voice stream into a VEM-Token1 sequence, convert the accompaniment stream into a VEM-Token2 sequence, and add them to a preprocessing library. S400: Use the deep learning to generate a lyrics score, a VEM-Token singing voice score, a VEM-Token accompaniment score, and a VEM-Token music score respectively.

2. The method according to claim 1, wherein The S100 includes: S110: Based on more than one of the modalities, decompose the emotions into the VEM classification and construct the VEM coordinate system, specifically including: S111: For the independent emotions, which include the compositions where the emotions are independent of each other and have no association, construct a one-way one-dimensional axis, with the lowest point of the emotion as the coordinate 0 point and the highest point of the emotion as the maximum coordinate point. S112: For the opposite emotion pairs, which include two compositions with opposite emotions to each other, construct a two-way one-dimensional axis. Among them, take the midpoint of the opposite emotion pair as the coordinate 0 point, the highest point of the positive emotion as the positive maximum coordinate point, and the highest point of the negative emotion as the negative maximum coordinate point. S113: For the associated opposite emotion groups, which include opposite emotion pairs with associations between more than one group of two opposite emotion pairs, align their respective coordinate 0 points, make their respective two-way one-dimensional axes super-orthogonal, divide them by a hyperplane, place their respective positive emotions on the same side of the hyperplane, and place their respective negative emotions on the opposite side of the hyperplane to construct a super-orthogonal coordinate system, with the highest point of their respective positive emotions as the positive maximum super-orthogonal coordinate point and the highest point of their respective negative emotions as the negative maximum super-orthogonal coordinate point; the vocal styles include one or a combination of ethnic song singing methods, pop song singing methods, Western song singing methods, popular song singing methods, original song singing methods, and opera singing methods.

3. The method according to claim 2, wherein The S100 also includes: S120: Construct the VEM function, including: S121: For the one - dimensional unidirectional coordinate system of the independent emotion, mark emotion scales on the one - dimensional unidirectional coordinate axis according to the attributes of the independent emotion, and construct the VEM function of the independent emotion; S122: For the two - dimensional unidirectional coordinate system of the opposite emotion pair, mark emotion scales on the two - dimensional unidirectional coordinate axis according to the attributes of the opposite emotion pair, and construct the VEM function of the opposite emotion pair; S123: For the hyper - orthogonal coordinate system of the associated opposite emotion group, construct the VEM function of the associated opposite emotion group according to the attributes between the opposite emotion pairs included in the associated opposite emotion group, and the VEM projection relationship between the emotion scales on one two - dimensional unidirectional coordinate axis and the emotion scales on another two - dimensional unidirectional coordinate axis; S124: The VEM function attributes include the numerical values of the emotion scales of the emotion on the coordinate axis, The VEM function relationship includes the calculation method between the numerical values of the emotion scales, and the calculation method includes one or a combination of the switching function, linear function, non - linear function, trigonometric function, and custom function between the emotions; The VEM projection relationship includes one or a combination of trigonometric functions and custom functions; S125: Add the VEM classification, the VEM coordinate system, and the VEM function to the VEM library.

4. The method according to claim 3, characterized in that The S200 includes: S210: Collect the vocal samples, where the vocal samples include song recordings of one or more of the VEM classifications and one or more of the vocal styles. Each vocal sample is evaluated by one or more human vocal experts, and the scores of the emotion judgment are given respectively for the singing and the accompaniment. The scores and the vocal samples are added to the VEM library; S220: The emotion judgment is obtained through steps including S221, S222, and S223, specifically including: S221: For the independent emotion, the scores are given respectively for the singing and the accompaniment by the vocal expert according to the coordinate from 0 point to the maximum point of the coordinate; S222: For the opposite emotion pair, the scores are given respectively for the singing and the accompaniment by the vocal expert according to the 0 point of the coordinate, the positive maximum point of the coordinate, and the negative maximum point of the coordinate; S223: For the associated opposite emotion group, the scores are given respectively for the singing and the accompaniment by the vocal expert according to the 0 point of the hyper - orthogonal coordinate system, the positive maximum point of the hyper - orthogonal coordinate, and the negative maximum point of the hyper - orthogonal coordinate; S230: According to the vocal samples and the scores in the VEM library, use supervised learning and deep learning to train the VEM function and the VEM parameters included in the VEM function, and add the scores and the VEM parameters to the VEM library.

5. The method according to claim 4, characterized in that, The S300 includes: The VEM processor performs hierarchical processing on the vocal file, and specifically further includes: L1.0 layer processing: Use the VEM processor to convert the vocal files of one or more and more than one channel into spectrum - format files of a unified format, where, If the vocal music file is an analog signal file, it is converted from analog to digital using a sampling frequency and converted into the spectral format file; If the vocal music file is a digital signal file, it is converted into the spectral format file; The sampling frequency includes 44.1 kHz, 48 kHz, or an integer multiple of 44.1 kHz and 48 kHz; L1.1 layer processing: If the spectral format file includes more than two channels, it is converted into the spectral format file of L1.1 layer mono; L1.1.1 layer processing: For the spectral format file of L1.1 layer mono, perform Fourier transform, constant Q transform, Mel spectrum transform, Hilbert transform, discrete wavelet transform, etc., to generate the L1.1.1 layer spectrogram; L1.1.1.1 layer processing: For the L1.1.1 layer spectrogram, set up rhythm pointers, which respectively point to the detected spectral energy mutation points, or percussion rhythm points, or dynamic envelope rhythm points, or their combinations, Align the rhythm pointers and pre-mark with the rhythm pointers as the beat starting points; L1.1.2 layer processing: For the complete L1.1 layer, use a convolutional recurrent neural network for the deep learning to circularly detect the rhythm pointers. Those with the pre-marked beat starting points coinciding are confirmed as the beat starting point marks, and those not coinciding are marked as variable rhythm starting point marks. Starting from the variable rhythm starting point marks, circularly execute the L1.1.1.1 layer processing to confirm the beat sequence of the complete vocal music file; L1.2 layer processing: According to the beat starting point marks, divide the spectral format file of the entire vocal music file into one VEM-Token according to the content of the spectral format file in each beat. According to the beat sequence, divide the entire spectral format file into a VEM-Token sequence and add it to the preprocessing library.

6. The method according to claim 5, wherein The S300 also includes: L1.4 layer processing: For the L1.1 layer and the L1.2 layer, set up beat intensity marks and beat type marks. Taking the VEM-Token as the unit, detect the spectral energy and divide it into pre-strong beats, pre-sub-strong beats, and pre-weak beats according to the spectral energy intensity; L1.4.1 layer processing: For the complete L1.1 layer, use a convolutional recurrent neural network and a Bayesian model to perform beat intensity detection and statistics, and circularly detect and mark the pre-strong marks, pre-sub-strong beats, and pre-weak marks; If the probability of the pre-sub-strong beat appearing is lower than the probability determination value, it is determined that the beat type mark of the vocal music file includes 2 / 4 beat or 3 / 4 beat or 3 / 8 beat, that is, strong, weak, and strong, weak, weak beat types; If the probability of the pre-sub-strong beat appearing is higher than the probability determination value, it is determined that the beat type mark of the vocal music file includes 4 / 4 beat, that is, strong, weak, sub-strong, weak beat type. The probability determination value is selected from 1% to 25%; During the loop detection, if the pre-strong mark, pre-secondary strong beat and pre-weak mark in the previous and next sequences overlap, it is confirmed as a continuous rhythm type; if they do not overlap, they are marked as a change-strong weak point, and the L1.4 layer processing is executed cyclically starting from the change-strong weak point to confirm the complete strong and weak sequence of the vocal file; L1.5 layer processing: for the vocal file, marking the beat type and measure mark, and adding it to the pre-processing library, the beat type includes at least one of 2 / 4 beat, 3 / 4 beat, 4 / 4 beat, and 3 / 8 beat; L1.6 layer processing: according to the L1.1 layer, the L1.2 layer and the L1.5 layer, the frequency of the VEM-Token is measured, and the output key signature is calculated according to the twelve-tone equal temperament and the rules of the simple notation and the five-line notation in music theory, and the deep learning is performed using the convolutional recurrent neural network to cyclically detect the key signature. If a key change is found, the new key signature and key change mark are recorded and added to the preprocessing library; L1.7 layer processing: according to the twelve equal temperament rules in music theory and the spectrum format file, for each VEM-Token, the note name sequence 1 present therein is calculated, for the VEM-Token sequence, the note name sequence 2 is calculated, for the entire vocal file, the note name sequence 3 is constructed and added to the preprocessing library; The note names are symbols used in music theory to record pitches using the twelve-tone equal temperament, and the note name sequence is a musical score that records the note names and durations.

7. The method according to claim 6, characterized in that, The S300 further includes the following mono processing steps: S330: For the monophonic vocal file, the VEM processor is used according to the spectrum format file, the preprocessing library and the VEM library content, and the deep learning method is used, specifically including the dual-stream U-Nnt, Demucs or Spleeter encoding and decoding architecture in the convolutional recurrent neural network. Under the synchronization constraint of the beat sequence, the encoder is subjected to multi-layer convolution downsampling, and the decoder is subjected to multi-layer convolution upsampling, and the vocal stream and the accompaniment stream are decomposed into the singing stream and the accompaniment stream, and the streams are added to the preprocessing library; S340: according to the beat sequence, decompose the singing stream into the VEM-Token1 and the VEM-Token1 sequence, decompose the accompaniment stream into the VEM-Token2 and the VEM-Token2 sequence, and add them to the preprocessing library; S350: The singing voice stream and the accompaniment stream of the complete spectrum format file are synthesized into an L2.0 layer accompaniment stream file and an L3.0 layer singing voice stream file, and added to the pre-processing library.

8. The method according to claim 7, wherein The S300 further includes the following multi-channel processing steps: S360: For the multi-channel vocal music file, the steps include S361, S362 and S363, specifically: S361: For the vocal music files including 2 channels, 4 channels, 5 channels and 5.1 channels, use the file content of each channel, convert it into the spectral format file, obtain the VEM-Token and the VEM-Token sequence of each channel, and determine the VEM-Token1 and the VEM-Token1 sequence, and the VEM-Token2 and the VEM-Token2 sequence according to the format of the vocal music file; S362: For the vocal music files other than 2 channels, 4 channels, 5 channels and 5.1 channels, use the file content of each channel, convert it into the spectral format file, obtain the VEM-Token and the VEM-Token sequence of each channel, and determine the VEM-Token1 and the VEM-Token1 sequence, and the VEM-Token2 and the VEM-Token2 sequence according to the format of the vocal music file; S363: Use a loss function for training to split and obtain the L4.0 layer singing stream file and the L5.0 layer accompaniment stream file.

9. The method according to claim 7 or 8, characterized in that, The specific steps of S400 include: S410: Use speech recognition to recognize the VEM-Token1 sequence, form the lyrics, align the lyrics and the beat start marker, and split to obtain the lyric scores of the L1.8 layer single sentence lyrics, the L1.9 layer single paragraph lyrics, and the L1.A layer full song lyrics; S420: According to the single sentence lyrics, single paragraph lyrics and full song lyrics, and based on the VEM library, use the convolutional recurrent neural network and the Bayesian model to respectively and circularly detect the VEM classification in the VEM-Token1 and the VEM-Token1 sequence, calculate the distribution probability, upgrade the VEM parameters, and add them to the preprocessing library; S430: For the complete vocal music file, according to the VEM classification and the distribution probability, align the lyrics and the beat start marker, calculate and output the VEM classification with the largest and the second largest probabilities in the distribution probability, and form the VEM-Token singing score.

10. The method according to claim 9, wherein S400 also includes: S440: Construct an instrument sound effect library and add it to the VEM library; S450: According to the VEM library, use the convolutional recurrent neural network and the Bayesian model to respectively and circularly detect the instrument sound effect library in the VEM-Token2 and the VEM-Token2 sequence, calculate the matching probability, and add it to the preprocessing library; S460: For the complete vocal music file, according to the VEM classification and the instrument matching probability, align the lyrics and the beat start marker, calculate and output the VEM classification with the largest probability in the matching probability, and calculate and obtain the VEM-Token accompaniment scores of the numbered musical notation and the staff notation according to the pitch name sequence 3 and the rules of music theory; S470: Align with the starting beat marker according to the VEM-Token accompaniment score, and merge the VEM-Token vocal score, the VEM-Token accompaniment score, and the lyrics score into the VEM-Token music score.

11. The method according to claim 1 or 10, characterized in that It further includes: Access the resources of natural language processing NLP or large language model LLM according to the lyrics or lyrics score to generate a vocal emotion score; Access the resources of natural language processing NLP or large language model LLM according to the accompaniment stream and the VEM-Token accompaniment score to generate an accompaniment emotion score; Synthesize the vocal emotion score and the accompaniment emotion score to obtain a synthesized music score; According to the communication protocol, connect the VEM-Token, the VEM library, and the preprocessing library to a common large model system to customize a vocal intelligent agent Agent.

12. The method according to claim 9, wherein It further includes: Access the resources of natural language processing NLP or large language model LLM according to the lyrics or lyrics score to generate a vocal emotion score; Access the resources of natural language processing NLP or large language model LLM according to the accompaniment stream and the VEM-Token accompaniment score to generate an accompaniment emotion score; Synthesize the vocal emotion score and the accompaniment emotion score to obtain a synthesized music score; According to the communication protocol, connect the VEM-Token, the VEM library, and the preprocessing library to a common large model system to customize a vocal intelligent agent Agent.

Citation Information

Patent Citations

  • Voice signal analysis sub-system based on multi-modal emotion identification system

    CN108899050A

  • Multi-modal emotion recognition model and method based on text and voice confidence

    CN115358212A