Multi-language vocal music and lyric alignment analysis system based on multi-dimensional fusion data analysis

By constructing a multilingual vocal and lyrics alignment analysis system, and utilizing multimodal pre-trained feature contrastive learning and cross-modal feature space models, the system solves the problem of integrating vocal data from different languages, achieves accurate identification of multilingual vocal production and intelligent assistance in music analysis, and improves teaching effectiveness.

CN120910299APending Publication Date: 2025-11-07CENTRAL CONSERVATORY OF MUSIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511017703.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate vocal data from different languages ​​to establish a precise correlation model between vocalization and phoneme symbols, resulting in an inability to accurately identify vocal characteristics and affecting music analysis and teaching effectiveness.

Method used

A multilingual vocal and lyric alignment analysis system based on multidimensional fusion data analysis is constructed. Vocal and lyric alignment is automatically performed through multimodal pre-trained feature contrast learning. Unified phoneme annotation, multi-threaded parallel audio slicing, improved MERT model and BERT architecture are used for feature extraction and alignment. A joint representation model across modal feature spaces is constructed, and alignment results are generated through an interactive feedback module.

Benefits of technology

It enables rapid analysis and accurate recognition of vocal production in multiple languages, improving the effectiveness and efficiency of music analysis and teaching, and providing intelligent auxiliary analysis solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910299A_ABST
    Figure CN120910299A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-language vocal music and lyric alignment analysis system based on multi-dimensional fusion data analysis, and the system comprises a data set construction module which is used for obtaining original multi-source data and processing the original multi-source data to obtain a structured multi-dimensional data set; the intelligent alignment module is used for automatically aligning vocal music and lyrics through multi-modal pre-training feature comparative learning based on the structured multi-dimensional data set to obtain a space-time alignment result; and the interaction feedback module is used for visually presenting vocal music lyric vocal characteristics through the space-time alignment result. Through the data-driven intelligent alignment module, an innovative intelligent auxiliary analysis scheme is provided for the fields of music analysis and music education, and the effect and efficiency of music analysis and music teaching are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent music analysis and intelligent music education, and particularly relates to a multi-language vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis. BACKGROUND

[0002] In the field of music analysis and intelligent technology fusion, it has become a crucial research direction to assist multi-language cross-regional vocal lyrics singing with the aid of technological means. This field not only concerns cultural heritage, but also enables music researchers to quickly analyze multi-language vocal singing through intelligent tools, which has significant academic value and application prospects. However, there is currently no open source method suitable for intelligent analysis of multi-language vocal sound to sound alignment.

[0003] Under this background, the field faces many challenges. The first and foremost is how to integrate vocal data from different languages to form a unified and accurate sound analysis framework. Since the vocal annotation of multi-language singing is often based on the phonetic system of its own language, it is limited to the Latin alphabet, which leads to the following two problems: (1) the same sound is represented by different symbols in different language systems; (2) the same symbol represents different sounds in different language systems. Therefore, it is impossible to establish an effective mapping from symbols to sounds. Another core problem is how to accurately identify the singing in the continuous singing of vocal and establish a correlation model between the singing and the phonetic symbols. Without such a correlation model, the system is difficult to judge the singing features and cannot provide singing information that is helpful for music analysis for researchers.

[0004] Therefore, how to construct a precise correlation model between singing and phonetic symbols based on the unified alignment of multi-language vocal phonetic data, and to realize the effective identification of vocal singing features accordingly, has become a key problem that needs to be solved by the present application. SUMMARY

[0005] To solve the above technical problems existing in the prior art, the present application proposes a multi-language vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis, which helps to improve the effect and efficiency of music analysis and music teaching.

[0006] To achieve the above purpose, the present application provides a multi-language vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis, comprising:

[0007] A data set construction module is used to obtain original multi-source data and process the original multi-source data to obtain a structured multi-dimensional data set.

[0008] An intelligent alignment module is configured to automatically align vocals and lyrics based on the structured multi-dimensional data set through multimodal pre-training feature contrast learning to obtain a spatiotemporal alignment result.

[0009] An interactive feedback module is configured to visually present vocal lyrics singing features through the spatiotemporal alignment result.

[0010] Preferably, the data set construction module comprises a unified phoneme annotation unit, a unified alignment format unit, and a parallel audio slicing unit.

[0011] The unified phoneme annotation unit is configured to unify phoneme annotations of the original multi-source data, construct a phonetic symbol-IPA key-value pair rule dictionary of different languages, and map phoneme symbols of different languages to IPA symbols.

[0012] The unified alignment format unit is configured to record start timestamp-end timestamp-phoneme symbol data, represent timestamps in seconds as floating-point numbers, and verify timestamp validity to obtain an aligned multi-source data set in a unified format.

[0013] The parallel audio slicing unit is configured to align and slice the aligned multi-source data set in a multi-thread parallel manner to obtain an aligned multi-source data set of each audio segment less than a first preset time length.

[0014] Preferably, the unified phoneme annotation unit comprises:

[0015] A symbol mapping engine is configured to convert multi-language phoneme symbols to IPA through traversal retrieval and replacement.

[0016] A verification subunit is configured to verify mapping accuracy by combining manual sampling with automatic testing.

[0017] Preferably, the unified alignment format unit comprises:

[0018] A timestamp verification subunit is configured to verify the validity of timestamps by setting verification standards, wherein the verification standards include: an end timestamp greater than a start timestamp, and a difference between a final end timestamp and an audio total duration less than a second preset time length.

[0019] An exception handling subunit is configured to manually review and automatically correct audio segments that exceed an error range.

[0020] Preferably, the intelligent alignment module comprises:

[0021] A multimodal pre-processing unit is configured to perform feature extraction and alignment preprocessing on audio and text respectively.

[0022] A contrast learning unit is configured to construct a joint representation model of a cross-modal feature space.

[0023] an optimizer unit configured to optimize the joint representation model of the cross-modal feature space by employing an optimization strategy for joint training of the CTC loss and the contrastive loss.

[0024] Preferably, the multi-modal pre-processing unit comprises:

[0025] an audio encoder configured to extract frequency domain features based on the improved MERT model;

[0026] a text encoder configured to construct a context-aware representation of the phoneme sequence based on the BERT architecture;

[0027] a feature alignment module configured to achieve the spatio-temporal alignment of the cross-modal features by the attention mechanism.

[0028] Preferably, the audio encoder comprises:

[0029] a spectrum conversion sub-unit configured to perform time-frequency conversion by employing the improved log-mel spectrum algorithm;

[0030] a resampling sub-unit configured to perform band-limited interpolation down-sampling for high sampling rate audio;

[0031] a padding alignment sub-unit configured to achieve the time sequence alignment processing of the intra-batch audio data.

[0032] Preferably, the text encoder comprises:

[0033] a word segmentation sub-unit configured to construct an international phonetic alphabet word segmenter based on the Unigram language model;

[0034] a sub-word encoding sub-unit configured to map the phoneme sequence into an integer encoding sequence;

[0035] a mask padding sub-unit configured to generate a text feature matrix with consistent length.

[0036] Preferably, the contrastive learning unit comprises:

[0037] a feature fusion sub-unit configured to calculate a similarity matrix of the audio and text features;

[0038] a loss function sub-unit configured to optimize the model by employing a weighted sum of the ForwardSum loss term and the CTC loss term;

[0039] an adversarial training sub-unit configured to introduce an adversarial sample generation mechanism to enhance the robustness of the model.

[0040] Preferably, the interaction feedback module comprises a feedback generation unit, wherein the feedback generation unit is configured to generate audio features and phoneme features based on an audio encoder and a text encoder, calculate an audio-to-phoneme alignment result for the audio features and the phoneme features respectively by using a DTW algorithm, and output to the user.

[0041] Compared with the prior art, the present application has the following advantages and technical effects:

[0042] The application discloses a multi-language vocal music and lyrics alignment analysis system based on multi-dimensional fusion data analysis, constructs a joint representation model of a cross-modal feature space based on a vocal phoneme alignment data set representing vocal music sound of lyrics by an international phonetic alphabet, continuously optimizes the model to improve the accuracy of feedback, and this method can help music researchers to analyze multi-language vocal sound faster and obtain effective research data. Through the data-driven intelligent alignment module, an innovative intelligent auxiliary analysis scheme is provided for musicology analysis and music education field, and the effect and efficiency of music analysis and music teaching are improved. BRIEF DESCRIPTION OF DRAWINGS

[0043] The accompanying drawings, which form a part of this application, are intended to provide further understanding of the application and are incorporated herein in their entirety, and the illustrative embodiments thereof and their description serve the purpose of explanations rather than limiting the present application. In the drawings:

[0044] Figure 1 A multi-language vocal music and lyrics alignment analysis system structure schematic diagram based on multi-dimensional fusion data analysis is provided for the embodiment of the present application.

[0045] Figure 2 A phoneme text pre-encoding flowchart is provided for the embodiment of the present application.

[0046] Figure 3 A text encoder processing process schematic diagram is provided for the embodiment of the present application.

[0047] Figure 4 A comparative learning unit processing process schematic diagram is provided for the embodiment of the present application.

[0048] Figure 5 A process schematic diagram for obtaining a comparative loss term is provided for the embodiment of the present application. DETAILED DESCRIPTION

[0049] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0050] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown.

[0051] As Figure 1 The embodiment proposes a multi-language vocal music and lyrics alignment analysis system based on multi-dimensional fusion data analysis, comprising:

[0052] A data set construction module is configured to obtain original multi-source data and process the original multi-source data to obtain a structured multi-dimensional data set.

[0053] An intelligent alignment module is configured to automatically align vocal music and lyrics based on the structured multi-dimensional data set through multi-modal pre-training feature comparison learning to obtain a spatio-temporal alignment result.

[0054] An interactive feedback module is configured to indicate vocal music and lyrics vocalization features through the spatio-temporal alignment result.

[0055] Further, the data set construction module comprises a unified phoneme annotation unit, a unified alignment format unit, and a parallel audio slicing unit.

[0056] The unified phoneme annotation unit is configured to unify phoneme annotations of the original multi-source data, construct a phonetic symbol-IPA key-value pair rule dictionary of different languages, and map phoneme symbols of different languages to IPA symbols.

[0057] The unified alignment format unit is configured to record start timestamp-end timestamp-phoneme symbol data, represent timestamps in seconds as floating-point numbers, and verify timestamp validity to obtain an aligned multi-source data set in a unified format.

[0058] The parallel audio slicing unit is configured to align and slice the aligned multi-source data set in a multi-thread parallel manner to obtain an aligned multi-source data set of each audio segment less than a first preset time length.

[0059] Specifically, the unified phoneme annotation unit comprises:

[0060] A symbol mapping engine is configured to convert multi-language phoneme symbols to IPA by traversing and searching for replacement;

[0061] A verification subunit is configured to verify mapping accuracy by combining manual sampling with automated testing.

[0062] The unified phoneme labeling unit is used for unified phoneme labeling of original multi-source data records, and a "phonetic symbol-IPA" key-value pair rule dictionary of each language is constructed. The phonetic symbols of each language are mapped to IPA symbols through an algorithm process of traversal, retrieval and replacement.

[0063] The unified alignment format unit includes:

[0064] The timestamp verification subunit is used for verifying the validity of the timestamp by setting a verification standard, wherein the verification standard includes: the end timestamp is greater than the start timestamp, and the difference between the final end timestamp and the total duration of the audio is less than a second preset time length.

[0065] The abnormality processing subunit is used for manual review and automatic correction of audio segments that exceed the error range.

[0066] The unified alignment format unit is used for recording in the lab file format, recording in the form of "start timestamp-end timestamp-phoneme symbol", using a second-level floating point number to represent the timestamp, verifying the validity of the timestamp, and obtaining a unified format of the aligned multi-source data set.

[0067] The parallel audio slicing unit uses a multi-thread parallel method to align and slice the unified format of the aligned multi-source data set, and uses the midpoint of the "silence" labeled timestamp as the slicing point to obtain an aligned multi-source data set with each audio segment being less than 30s.

[0068] In the unified alignment format, a second-level floating point number is used to represent the timestamp, considering the high precision and stable storage of the floating point number.

[0069] The validity of the timestamp is verified by the following standards:

[0070] A. The end timestamp is greater than the start timestamp.

[0071] B. The difference between the final end timestamp and the total duration of the audio is less than 10ns.

[0072] The data that does not meet the validity is screened out, and the abnormal reason is checked by sampling. For audio with a long tail silence segment, the phonetic text label is cut to obtain a unified format of the aligned multi-source data set.

[0073] As a specific embodiment of the present embodiment, the multi-thread parallel algorithm follows the following logic: (pseudo code)

[0074] Define global parameters:

[0075] MAX_WORKERS = 4 / / maximum number of worker threads per thread pool

[0076] NUM_THREAD_POOLS = 3 / / number of thread pools used

[0077] TASKS = [... ] / / list of all tasks to process

[0078] Initialize multiple thread pools:

[0079] THREAD_POOLS = [ThreadPoolExecutor(MAX_WORKERS) for in range(NUM_THREAD_POOLS)]

[0080] Distribute tasks to each thread pool:

[0081] BATCH_SIZE = len(TASKS) / NUM_THREAD_POOLS

[0082] batches = split_into_batches(TASKS, BATCH_SIZE)

[0083] For each batch and corresponding thread pool:

[0084]

[0085] Shutdown all thread pools:

[0086] for pool in THREAD_POOLS:

[0087] pool.shutdown(wait=True)

[0088] Further, the intelligent alignment module comprises:

[0089] A multi-modal preprocessing unit for feature extraction and alignment preprocessing of audio and text respectively;

[0090] A contrastive learning unit for constructing a joint representation model of cross-modal feature space;

[0091] An optimizer unit for optimizing the joint representation model of the cross-modal feature space using an optimization strategy of joint training of CTC loss and contrastive loss.

[0092] Specifically, the multi-modal preprocessing unit comprises:

[0093] An audio encoder for extracting frequency domain features based on an improved MERT model;

[0094] A text encoder for constructing a context-aware representation of phoneme sequences based on a BERT architecture;

[0095] The feature alignment module is configured to realize the spatio-temporal alignment of the cross-modal features through an attention mechanism.

[0096] The audio data in the aligned multi-source data set is converted from time domain to frequency domain through an improved log-mel spectrum algorithm for audio preprocessing, and the audio data in a batch is filled and aligned to obtain a processed audio data batch.

[0097] The phoneme text data in the aligned multi-source data set is segmented through an Unigram algorithm for phoneme text preprocessing, and the phoneme text data in a batch is filled and aligned to obtain a processed phoneme text data batch.

[0098] The contrast learning unit is configured to train the phoneme encoding batch and the audio encoding batch through a contrast learning strategy to obtain an audio encoder and a phoneme text encoder suitable for vocal song lyric phoneme alignment, and construct a joint representation model of a cross-modal feature space.

[0099] The audio encoder comprises:

[0100] The spectrum conversion subunit is configured to perform time-frequency conversion through an improved log-mel spectrum algorithm.

[0101] Specifically, the audio data in the aligned multi-source data set is converted from time domain to frequency domain through an improved log-mel spectrum algorithm for audio preprocessing, and the audio data in a batch is filled and aligned to obtain a processed audio data batch.

[0102] The resampling subunit is configured to perform band-limited interpolation downsampling based on fast Fourier transform for high sampling rate audio.

[0103] The padding alignment subunit is configured to realize the time sequence alignment processing of the audio data in a batch.

[0104] In the spectrum conversion subunit, the process of time-frequency conversion through the improved log-mel spectrum algorithm includes: performing framing and windowing preprocessing on the audio signal, performing short-time Fourier transform (STFT) and power calculation to obtain a power spectrum, multiplying the power spectrum by a mel filter bank to obtain a mel spectrum, wherein the formula for converting linear frequency to mel frequency is as follows:

[0105]

[0106] In the formula, f is the linear frequency (unit: Hz), m is the Mel frequency, and the formula for calculating the Mel spectrum is:

[0107] M[m]=∑ k H m (f k )·|X[k]| 2 ,

[0108] In the formula, M[m] is the spectral energy of the m-th Mel frequency channel. H m (f k ) is the m-th Mel filter at frequency f k The response value at (typically a triangular filter). |X[k]| 2 It is the spectrum at frequency f k The power (energy) at that point. Taking the logarithm of the Mel spectrum to enhance the dynamic range yields the audio logarithmic Mel spectrum data.

[0109] When the audio sampling rate is greater than 16kHz, the resampling subunit will trigger a high-quality band-limited interpolation resampling mechanism based on fast Fourier transform to obtain audio downsampled to 16kHz.

[0110] Based on the Whittaker-Shannon interpolation formula:

[0111]

[0112] In the formula, x[n] is the original sampling point, T is the original sampling period, and t is the new time point.

[0113] Based on the ratio of the target sampling rate to the original sampling rate, an FIR filter is designed to avoid frequency aliasing. The filter is decomposed into multiple sub-filters, each corresponding to a different phase shift. A window function is used to truncate the sinc function, forming a finite-length interpolation kernel.

[0114] h k (n) = w(n)sinc(nk),

[0115] In the formula, h k w(n) is the filter coefficient corresponding to the k-th phase, w(n) is the window function, and sinc(nk) is the interpolation basis function.

[0116] The interpolated output y[m] is obtained:

[0117]

[0118] In the formula, m is the m-th sample of the output signal, and h k It is the filter corresponding to the current interpolation phase k, and N is the interpolation kernel length, which is 512.

[0119] Further, the text encoder comprises:

[0120] A word segmentation subunit for constructing an international phonetic alphabet word segmenter based on a Unigram language model;

[0121] A subword encoding subunit for mapping a phoneme sequence into an integer encoding sequence;

[0122] A mask padding subunit for generating a text feature matrix with consistent length.

[0123] The word segmentation subunit for constructing an international phonetic alphabet word segmenter based on a Unigram language model comprises: constructing a word segmenter vocabulary based on a phoneme text corpus in a database through a Unigram language model algorithm, constructing a seed vocabulary from the corpus, and the size of the seed vocabulary is twice the size of the target vocabulary, and finally the target vocabulary and the corresponding subword probability p(x i ) are estimated by maximizing the following likelihood function on the training corpus:

[0124]

[0125] wherein, is the entire training corpus, is the number of sentences in the corpus, P(x) is the probability of the segmentation x of the sentence X (s) , and is the set of all possible segmentations of the sentence X (s) .

[0126] The expectation maximization (EM) algorithm is used to optimize

[0127] Based on the current vocabulary and the probability p(x i ), the expected number of occurrences E(x) of each subword x in the corpus is calculated;

[0128] The subword probability is updated using the expected count:

[0129] All low-probability subwords satisfying p(x) < θ (where θ is a preset threshold) are deleted, or the bottom-ranked (e.g., the lowest 10%) subwords in the vocabulary are deleted;

[0130] If the current vocabulary size is still greater than the target size, repeat the above steps to finally obtain a word segmenter trained based on an international phonetic alphabet corpus.

[0131] The phoneme text data in the aligned multi-source data set is segmented into subwords by a word segmenter, and then mapped to an integer vector (one phoneme symbol corresponds to one scalar) by a vocabulary, and the length of the subword vector is recorded. All subword vectors in a sample are spliced to obtain a spliced subword vector corresponding to a sample, and the spliced subword vectors of all samples in a batch are combined to obtain a phoneme text batch.

[0132] A word vector length mask matrix is generated according to the subword vector length of the phoneme text of a sample in the batch, the number of rows of the length mask matrix is the total number of subwords, and each row sums to 1 on the corresponding subword length.

[0133] Further, according to the model training requirements, the tail alignment padding is performed on the audio log mel spectrum data and the phoneme text segmentation data of different lengths in a batch, so that the lengths of all data are less than or equal to the maximum length set by the system, the maximum data length in the batch is taken as the padding target length, the tail of the audio log mel spectrum data is padded with a floating point number 0.0, and the tail of the phoneme text segmentation data is padded with an integer 0.

[0134] An audio data batch matrix is obtained, the number of rows of which is the number of samples in a batch, and each row is the padded sample audio log mel spectrum data.

[0135] An audio data batch matrix is obtained, the number of rows of which is the number of samples in a batch, and each row is the padded sample audio log mel spectrum data.

[0136] After padding the data, an attention mask matrix is generated, the number of rows of which is the number of samples in a batch, and the length of each column is the padding target length, the matrix element is a floating point number 1.0 at the position corresponding to the valid data of each sample, and the matrix element is a floating point number 0.0 at the position corresponding to the padded data.

[0137] After completing the preprocessing of the audio data and the phoneme text data, the pre-encoding stage is entered.

[0138] In the audio encoder, the audio data batch matrix and the corresponding audio attention mask matrix are input into the MERT model to obtain a 4-dimensional hidden state matrix with a feature dimension of 1024 in 25 layers, the 25-layer hidden state matrix is aggregated into a one-layer hidden state matrix through a 3-dimensional convolution kernel, and then a linear projection layer is used to obtain an audio feature matrix after normalization.

[0139] In the text encoder, as Figure 3As shown, the phoneme text data batch matrix and the corresponding phoneme text attention mask matrix are input into the BERT model. The phoneme text data batch matrix is input into a word embedding layer to obtain a subword embedding matrix, input into a position embedding layer to obtain a position embedding matrix, and input into a segment embedding layer to obtain a segment embedding matrix. The phoneme text data batch matrix, the position embedding matrix, and the segment embedding matrix are summed to obtain a sum embedding matrix, and the sum embedding matrix is input into a layer normalization layer to obtain an output embedding matrix. After the embedding part is completed, the encoding part is entered.

[0140] The encoding part is composed of a plurality of transformer encoders connected in series. Each transformer encoder performs the following steps, as shown in the following formula: Figure 2

[0141] S1, the phoneme text attention mask matrix is expanded, and a multi-head attention mask matrix is calculated based on the output embedding matrix to meet different attention requirements;

[0142] S2, the output embedding matrix and the multi-head attention mask matrix are connected in residual and normalized by a layer normalization layer;

[0143] S3, the output obtained in S2 is input into a feedforward neural network sublayer;

[0144] S4, the output obtained in S3 is connected in residual and normalized by a layer normalization layer to obtain the final transformer encoder output.

[0145] The final output of the multi-layer encoder is input into a linear projection layer and a mean pooling layer to obtain a final output. The final output is multiplied by a word vector length mask matrix, and a phoneme feature matrix is obtained after normalization.

[0146] In this embodiment, the linear projection layer in the audio encoder and the linear projection layer in the text encoder have the same output dimension.

[0147] Further, the contrastive learning unit includes:

[0148] A feature fusion subunit is configured to calculate a similarity matrix of the audio feature matrix and the phoneme feature matrix.

[0149] A loss function subunit is configured to use a ForwardSum+CTC composite loss function to optimize the model.

[0150] An adversarial training subunit is configured to introduce an adversarial sample generation mechanism to enhance the robustness of the model.

[0151] Specifically, as shown in the following formula: Figure 4 The similarity matrix SIM is calculated based on the audio feature matrix and the phoneme feature matrix.

[0152] ​

[0153] where S is the audio feature matrix, P is the phoneme feature matrix, and t is the temperature parameter.

[0154] The loss function is calculated via the following steps:

[0155] S1. Add a blank symbol in the second dimension of the similarity matrix, and then traverse a batch and operate on each sample in it;

[0156] S2. Construct the target sequence:

[0157] Target = (1, 2, …, L),

[0158] where L is the effective length of this sample in the phoneme text data batch matrix.

[0159] S3. Cut the similarity matrix to the actual audio data length and phoneme text data length of the current sample, and calculate the log_softmax to construct the probability matrix:

[0160] Z = log_softmax(SIM).

[0161] where SIM is the similarity matrix.

[0162] S4. Pass the probability matrix and the target sequence through the Connectionist Temporal Classification Loss (CTC Loss) algorithm to obtain the probability from the audio sequence of this sample to the target sequence, i.e., the sample CTC loss term;

[0163] S5. Sum and average the sample CTC loss terms obtained by calculating each sample in the batch to obtain the batch ForwardSum loss term;

[0164] S6. Calculate the softmax of the audio feature matrix after passing through the linear projection layer with an output dimension of the dictionary length to obtain the probability matrix from the audio to the full dictionary, and pass the probability matrix and the phoneme feature matrix (as the target sequence) through the CTC Loss algorithm to obtain the CTC loss term;

[0165] S7. Weighted sum the ForwardSum loss term and the CTC loss term to obtain the final loss Loss:

[0166] Loss = L ForwardSum + 0.1 * L ctc ,

[0167] The audio encoder and the phoneme text encoder parameters are updated by backpropagation of the total contrastive loss function for training.

[0168] Furthermore, during the phoneme text preprocessing, a random negative sample sequence is generated for each sample through the following steps:

[0169] S1. Calculate the length s of the character sequence that needs to be modified in the sample:

[0170]

[0171] In the formula, L is the total length of the sample sequence, and p is a pre-specified probability value in the interval [0,1).

[0172] S2. Generate a uniformly distributed random number array U = {u1, u2, ..., u...} of size s. s Given that each number has a value in the interval [0, 1), map these random numbers to their corresponding indices in the sequence:

[0173]

[0174] In the formula, I is a set of unique position indices;

[0175] S3. Replace, insert, or delete the characters corresponding to the position indices in I to obtain the negative sample sequence.

[0176] The negative samples are generated for all samples in a batch and then filled and integrated to obtain the phoneme text negative sample batch.

[0177] Optionally, the negative sample phoneme feature matrix is ​​obtained by encoding the negative sample phoneme text batches.

[0178] Furthermore, such as Figure 5 As shown, the contrastive loss function can be calculated in the contrastive learning unit, and the positive sample loss term L is calculated from the audio feature matrix and the phoneme feature matrix. pos :

[0179]

[0180] In the formula, S is the audio feature matrix, P is the phoneme feature matrix, t is the temperature parameter, and b is the bias parameter. S is the L2-normalized vector of the i-th row of the audio feature matrix. (i) Let |S| be the vector of the i-th row of the audio feature matrix. (i) ||2 is the length of the vector in the i-th row of the audio feature matrix. P is the L2-normalized vector of the j-th row of the phoneme feature matrix. (j) Let |P| be the vector in the j-th row of the phoneme feature matrix. (j) ||2 is the length of the vector in the j-th row of the phoneme feature matrix, Logits P Z is the output logits matrix. sis the L2 normalized audio feature matrix, is the transpose of the L2 normalized phoneme feature matrix.

[0181] The negative sample loss term L is calculated from the audio feature matrix and the negative sample phoneme feature matrix neg :

[0182]

[0183] In the formula, S is the audio feature matrix, N is the negative sample phoneme feature matrix, t is the temperature parameter, b is the bias parameter, is the L2 normalized vector of the i-th row of the audio feature matrix, is the L2 normalized vector of the i-th row of the negative sample phoneme feature matrix, (j) is the vector of the j-th row of the negative sample phoneme feature matrix, (j) is the length of the vector of the j-th row of the negative sample phoneme feature matrix, N is the output logits matrix, is the transpose of the L2 normalized negative sample phoneme feature matrix.

[0184] The total contrastive loss function is obtained by summing the positive sample loss term and the negative sample loss term. The audio encoder and the text encoder parameters are updated by backpropagation through the total contrastive loss function for training.

[0185] Further, the interactive feedback module includes a feedback generation unit, wherein the feedback generation unit is configured to generate audio features and phoneme features based on the audio encoder and the text encoder, calculate the audio-to-phoneme alignment result for the audio features and the phoneme features respectively using the DTW algorithm, and output to the user.

[0186] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in the present application can be easily thought of by those skilled in the art, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-lingual vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis, characterized in that, The method comprises the following steps: A dataset construction module is used to obtain original multi-source data and process the original multi-source data to obtain structured multi-dimensional data sets; An intelligent alignment module is used to automatically align vocals and lyrics based on the structured multi-dimensional data sets through multi-modal pre-training feature comparison learning to obtain a time-space alignment result; An interactive feedback module is used to visually present vocal and lyric singing features through the time-space alignment result.

2. The multi-lingual vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis according to claim 1, wherein, The dataset construction module comprises a unified phoneme annotation unit, a unified alignment format unit, and a parallel audio slicing unit; The unified phoneme annotation unit is used to unify the phoneme annotation of the original multi-source data, construct a phonetic symbol-IPA key-value pair rule dictionary of different languages, and map the phoneme symbols of different languages into IPA symbols; The unified alignment format unit is used to record start timestamp-end timestamp-phoneme symbol data, represent the timestamps in seconds, and verify the validity of the timestamps to obtain aligned multi-source data sets in a unified format; The parallel audio slicing unit is used to align and slice the aligned multi-source data sets in a multi-thread parallel manner to obtain aligned multi-source data sets with each audio segment being less than a first preset time length.

3. The multi-lingual vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis according to claim 2, wherein, The unified phoneme annotation unit comprises: A symbol mapping engine is used to convert multi-language phoneme symbols into IPA by traversing and searching for replacement; A verification subunit is used to verify the mapping accuracy by combining manual sampling with automatic testing.

4. The multi-lingual vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis according to claim 2, wherein, The unified alignment format unit comprises: A timestamp verification subunit is used to verify the validity of the timestamps by setting verification standards, wherein the verification standards include: the end timestamp is greater than the start timestamp, and the difference between the final end timestamp and the total audio duration is less than a second preset time length; An exception handling subunit is used to manually review and automatically correct audio segments that exceed the error range.

5. The multi-lingual vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis as claimed in claim 1, wherein, The intelligent alignment module comprises: A multi-modal pre-processing unit is used to extract and pre-process features of audio and text respectively; A contrastive learning unit is used to construct a joint representation model of a cross-modal feature space; An optimizer unit is used to optimize the joint representation model of the cross-modal feature space using a CTC loss and a contrastive loss joint training optimization strategy.

6. The multi-lingual vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis according to claim 5, wherein, The multi-modal pre-processing unit comprises: An audio encoder is used to extract frequency domain features based on an improved MERT model; A text encoder is used to construct a context-aware representation of phoneme sequences based on a BERT architecture; A feature alignment module is used to achieve time-space alignment of cross-modal features through an attention mechanism.

7. The multi-lingual vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis according to claim 6, wherein, The audio encoder comprises: A spectrum conversion subunit is used to perform time-frequency conversion using an improved log-mel spectrum algorithm; A resampling subunit is used to perform band-limited interpolation down-sampling for high sampling rate audio; A padding alignment subunit is used to perform time alignment processing of audio data within a batch.

8. The multi-lingual vocal and lyrics alignment analysis system based on multi- dimensional fusion data analysis according to claim 6, wherein, The text encoder comprises: A word segmentation subunit is used to construct an IPA word segmentation tool based on a Unigram language model; A subword encoding subunit is used to map phoneme sequences into integer encoding sequences; A mask padding subunit is used to generate a text feature matrix with consistent length.

9. The multi-lingual vocal and lyrics alignment analysis system based on multi-dimensional fusion data analysis as claimed in claim 5, wherein, The contrast learning unit comprises: a feature fusion subunit configured to calculate a similarity matrix of audio and text features; a loss function subunit configured to optimize the model by weighted summation of a ForwardSum loss term and a CTC loss term; an adversarial training subunit configured to introduce an adversarial sample generation mechanism to enhance the robustness of the model.

10. The multi-lingual vocal and lyrics alignment analysis system based on multi- dimensional fusion data analysis as claimed in claim 1, wherein, The interaction feedback module comprises a feedback generation unit, wherein the feedback generation unit is configured to generate audio features and phoneme features based on an audio encoder and a text encoder, calculate an audio-to-phoneme alignment result for the audio features and the phoneme features respectively by using a DTW algorithm, and output the audio-to-phoneme alignment result to a user.