Chinese song play cavity evaluation method based on large-scale audio language model
By constructing a multi-dimensional evaluation dataset of Chinese opera singing styles and a large-scale audio language model, combined with a music score encoder and feature connectors, the problem of the lack of interpretability in existing evaluation models is solved. This enables the simultaneous generation of scores and comments, improving the accuracy and interpretability of the evaluation and providing a unified and fair evaluation tool for opera singing style teaching and synthesis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAMEN UNIV
- Filing Date
- 2026-03-04
- Publication Date
- 2026-05-15
AI Technical Summary
Existing Chinese opera singing evaluation models only provide MOS scores, which lack interpretability and are devoid of text-based evaluation datasets and reliable, open, and fair evaluation methods. They also cannot provide comments and scores simultaneously.
A dataset for evaluating the singing style of Chinese Gezai Opera containing text comments was constructed. Based on the large-scale audio language model SALMONN, a music score encoder and feature connectors were added. The model was fine-tuned using low-rank adaptation techniques and trained using a combination of cross-entropy and mean squared error loss functions to generate comment text and total evaluation scores.
It enables the simultaneous generation of scores and comments during the evaluation process. The comment text provides detailed descriptions from multiple dimensions such as pitch and rhythm, which improves the interpretability and accuracy of the evaluation, fills the gaps in the dataset, and provides a unified and fair automated evaluation tool.
Smart Images

Figure CN122050433A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of audio signal processing and large-scale audio language model post-training technology, and particularly to the field of Chinese opera singing evaluation, specifically to a method for evaluating Chinese opera singing based on a large-scale audio language model. Background Technology
[0002] Artificial intelligence technology has made progress in the field of multimedia content generation, with rapid development and application of tasks such as text, audio, image, and video generation. Chinese opera, an intangible cultural heritage and national treasure, offers a new solution for the innovative development of Chinese opera through the synthesis of opera vocals. The input to the opera vocal synthesis task is lyrics, musical score, and speaker information; the output is synthesized opera vocal audio, which consists only of dry vocals without accompaniment.
[0003] In terms of data, the paper "Creating an A Cappella Singing AudioDataset for Automatic Jingju Singing Evaluation Research" by Gong Rong et al., published in the 4th International Digital Libraries for Musicology workshop (DLfM 2017), constructs an a cappella Jingju dataset containing 120 arias and 1265 melodic phrases for automatic Jingju singing evaluation research. In 2020, Wu Yusong et al., based on the aforementioned Peking Opera data and after post-processing, conducted research on the deviations in pitch and rhythm between actual Peking Opera performances and standard scores. The improvisation and personalized expression of Peking Opera actors can cause the rhythm and pitch contours of the actual performance to deviate significantly from the score. Based on the duration-aware attention network framework, they constructed an expressive vocal synthesis model: by using the Lagrange multiplier method to optimize the phoneme duration sequence under the constraint of the score note duration, the rhythm mismatch problem can be alleviated; for pitch deviation, pseudo scores generated from actual performances are used as training inputs, rather than directly from the original score. In 2023, Zhou Xun et al. published "A High-Quality Melody-Aware Peking Opera Synthesizer Using Data Augmentation" in the 2023 IEEE International Conference on Multimedia and Expo (ICME2023, pp. 1092-1097). This paper proposed the OperaSinger model for Peking Opera. The model uses FastSpeech2 as the backbone network and effectively improves the problem of poor naturalness of synthesized Peking Opera singing audio caused by ignoring local features by stacking melody-aware position-variable convolutional blocks and feedforward Transformer blocks in parallel in its encoder.In 2023, Bai Peng et al. published "Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local Modeling" in the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP2023, pp. 3302-3312). Addressing the problem of insufficient local modeling in the Transformer model, they proposed a local attention mechanism for neighboring phonemes and a phoneme-level adaptive weighted loss function, achieving improvements in Chinese pop song and Hokkien Gezi opera tasks. In 2024, Zheng Meizhen et al. published "FT-GAN: Fine-Grained Tune Modeling for Chinese Opera Synthesis" in *Thirty-Eighth Conference on Artificial Intelligence* (AAAI 2024, pp. 19697-19705), proposing the FT-GAN model for Chinese opera vocal synthesis, focusing on the problem of fine-grained melody modeling in vocal synthesis. Its core innovation lies in the construction stage of the Minnan Gezai Opera dataset: using a self-developed pitch extraction algorithm, it performs fine-grained annotation for different pitch variations of the same phoneme; in terms of model structure, FT-GAN is built on a generative adversarial network, and the discriminator achieves high-quality opera vocal synthesis through a complex design of a block-based discrimination mechanism. In 2025, Qiu et al., in their paper "HuangmeiSinger: A Dataset and A Branchformer-Diffusion Model for Huangmei Opera Synthesis" published in *The 4th International Conference on Computer, Artificial Intelligence and Control Engineering* (pp. 641–647), proposed a Huangmei Opera singing synthesis model, extending the research scope to Huangmei Opera. This model introduces a Branchformer encoder and a pitch diffusion module, specifically designed to capture and process the complex and diverse melodic features of Huangmei Opera, enhancing the expressiveness and adaptability of the synthesized singing style.In 2025, Li Yue et al. published "CODS: Audio-Text Alignment Dataset for Cantonese Opera Vocal Synthesis" in the Journal of South China University of Technology (Natural Science Edition) (pp. 1–10). They proposed a Cantonese opera vocal synthesis dataset with phoneme-level annotation and audio-text alignment. On this dataset, they conducted experiments using deep learning methods to achieve controllable Cantonese opera vocal synthesis with lyrics, timbre, and melody.
[0004] However, a significant challenge in the development of traditional Chinese opera vocal synthesis models is how to automatically evaluate synthesized audio or authentic vocal recordings. Existing research uses a combination of objective and subjective evaluation methods. Objective evaluation typically uses Voicing Decision Error (VDE), Gross Pitch Error (GPE), F0 Frame Error (FFE), and Logarithmic rooted mean square error of the fundamental frequency (F0 RMSE) to assess the fundamental frequency (F0) trajectory of synthesized traditional Chinese opera vocal audio. VDE measures the frame-level error rate of the real and synthesized audio in the "voiced frame" decision. GPE measures the frame-level error rate where there is a significant deviation between the synthesized F0 and the real F0 in frames where the "voicing decision is correct." FFE is a comprehensive indicator that measures the overall frame-level error rate of the synthesized audio's F0 trajectory; it integrates VDE and GPE to comprehensively reflect the quality of the F0 trajectory. All three indicators are measured as percentages. F0 RMSE is calculated by taking the logarithm of the fundamental frequency value and then calculating the root mean square error; the smaller the value, the more accurate the pitch trend of the synthesized vocals. Objective evaluation also uses Mel-cepstral distortion (MCD) to assess the similarity of the spectral characteristics between the synthesized audio and the real reference audio. However, all of the above evaluation metrics are reference-based, requiring a reference audio recording to calculate the results. Furthermore, these metrics only reflect one aspect, and the results for a particular aspect often differ from overall human perception.
[0005] In subjective evaluation, existing research uses the Mean Opinion Score (MOS) as a subjective evaluation indicator. MOS measures human subjective perception of the naturalness of audio, with a scoring range of 1 to 5. Evaluators are unaware of the audio source, ensuring the fairness and reliability of the results. In the evaluation of opera singing, since subjective evaluation is based on genuine human feelings, the focus of evaluation in this field is on subjective evaluation results. Different studies use different subjective evaluation conditions, and the subjective evaluation results of different models cannot be reproduced, lacking a transparent and fair automated evaluation method. To address this issue, in 2026, Bai Peng et al. published "Reference-free singing voice MOS prediction via multi-feature fusion, with integrated feature analysis" in *Applied Acoustics* (Vol. 241), proposing the Gezai opera singing MOS prediction dataset GOSMOS and proposing a multi-feature fusion no-reference MOS prediction model integrating initial perception, articulation clarity, timbre, pitch, and emotion.
[0006] While subjective MOS (Mutually Modular Structure) automatic prediction can alleviate this problem, it lacks interpretability and cannot clearly explain the core reasons for the singing score. Specifically, even if a MOS result is predicted for an audio file, the user still doesn't know exactly where the audio's flaws lie or how to improve it, hindering targeted improvements. Therefore, the evaluation of synthesized opera singing urgently needs a model that can provide both scores and descriptive text, especially for singing evaluations lacking sufficient textual descriptions of shortcomings. If a model could provide scores along with textual explanations for the evaluated audio, it would have significant application value in opera singing teaching, synthesis, and related tasks. Large-scale audio language models are multimodal models trained on massive amounts of audio-text pairing data. They can deeply integrate audio acoustic features and linguistic semantic information, enabling refined understanding and analysis of audio content.
[0007] In summary, existing technologies suffer from three main problems: First, current Chinese opera singing assessments only provide MOS scores without comments, thus lacking interpretability. Second, there is a lack of datasets for Chinese opera singing assessment research that include both comments and scores. Third, this field urgently needs a reliable, open, and fair large-scale assessment model that can accept both audio and sheet music inputs. Summary of the Invention
[0008] The purpose of this invention is to address the problem that existing Chinese opera singing evaluation models only provide scores and not comments. By constructing a high-quality professional dataset and fine-tuning a large-scale audio language model, this invention enables the simultaneous provision of comments and scores in Chinese opera singing evaluation.
[0009] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0010] A method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model includes the following steps:
[0011] Step 1: Construct a Chinese Gezai Opera singing evaluation dataset containing text comments. The dataset includes sentence-level dry vocal audio, corresponding musical scores, multi-dimensional evaluation text descriptions, and total singing evaluation scores. The audio sources include performances by professional singers, amateur singers, and synthesized audio generated by a Gezai Opera singing synthesis model. The multi-dimensional evaluation text descriptions are annotated by professional and public reviewers, and the descriptions provide comments and scores on aspects such as pitch, rhythm, range, timbre, singing techniques, volume variation, breath control, and emotional expression. A large text model is used to detect and refine the comments, preserving the original meaning and enhancing the diversity of textual expression. The dataset is then divided into training, validation, and test sets according to a set ratio.
[0012] Step 2: Construct a Gezai Opera singing evaluation model based on a large-scale audio language model: Based on the general large-scale audio language model SALMONN, add a music score encoder and a feature connector; the model input is the dry vocal audio to be evaluated and the corresponding music score; the music score information is used to evaluate basic normative evaluation indicators such as pitch and rhythm.
[0013] Step 3: Fine-tuning the Gezai Opera Singing Evaluation Model: The model input consists of the dry vocal audio and its corresponding score from the data sample constructed in Step 1. The model output is the predicted comment text and the total evaluation score. Fine-tuning employs Low-Rank Adaptation (LoRA) technology. During the model fine-tuning process, the model is constrained to generate comment text step by step according to the preset 8 evaluation dimensions. The text words of the comment text are trained under supervision using the cross-entropy loss function. For the word part of the total evaluation score output by the model, the mean squared error loss function is added for joint constraint.
[0014] Step 4: Evaluation and testing of Chinese Gezai Opera singing style: The trained Gezai Opera singing style evaluation model is applied to the automated evaluation of Gezai Opera singing audio, and index tests are performed on the test set divided in Step 1 to complete the performance verification of the model evaluation.
[0015] In step 1, the specific steps for constructing a Chinese Gezai Opera singing evaluation dataset containing text comments can be as follows:
[0016] 1.1 Data Acquisition and Synthesis: Collect authentic dry audio recordings of Taiwanese opera singing from publicly available datasets of synthesized Taiwanese opera singing; synthesize a preset number of dry audio recordings of Taiwanese opera singing based on a pre-trained Taiwanese opera singing synthesis model, ensuring that the audio sources cover professional singing, amateur singing, and synthesized audio, and achieving sample diversity.
[0017] 1.2 Score Matching: Collect the standard score corresponding to each dry vocal audio recording to ensure that the melody and rhythm of the score and audio are accurately matched. The score is in MIDI format.
[0018] 1.3 Standardized Crowdsourced Labeling: For each dry vocal recording, a standardized crowdsourced evaluation process is adopted, with professional judges with backgrounds in opera singing or opera teaching, as well as amateur judges, participating in the labeling; the ratio of professional judges to amateur judges is predetermined;
[0019] 1.4 Multi-dimensional evaluation: Each judge will write comments on the audio based on eight preset evaluation dimensions: pitch, rhythm, vocal range, timbre, singing technique, volume variation, breath control, and emotional expression, and will give a total evaluation score of 1-5.
[0020] The total evaluation score for each audio file is collected using a standardized crowdsourced evaluation process and a unified 1-5 point scoring standard for Taiwanese opera singing. The scoring standard can be as follows: 1 point indicates serious pitch and rhythm errors in the singing, making it impossible to hear properly; 2 points indicate multiple obvious errors in the singing, with poor overall expression; 3 points indicate the singing is basically complete, with a few acceptable minor errors; 4 points indicate the singing is fluent and accurate, with no obvious flaws; 5 points indicate the singing is perfect with no obvious flaws, at a professional level. The scoring interval can be 0.5 points.
[0021] 1.5 Optimization of Comment Text: In order to ensure the professionalism and diversity of comment text, a large text model is used to detect and polish the manually annotated comment text, generating comment text with diversified text expressions that do not change the original meaning.
[0022] 1.6 Dataset partitioning: Divide the optimized dataset into training, validation and test sets according to a preset ratio to ensure that the data in each subset does not overlap; the ratio can be 8:1:1.
[0023] In step 2, the specific steps for constructing the evaluation model of Gezai Opera singing style based on a large-scale audio language model can be as follows:
[0024] 2.1 Model Foundation Construction: Based on the general large-scale audio language model SALMONN, a music score encoder and a feature connector are added to construct a Gezai Opera singing evaluation model that supports dual-modal input of dry vocal audio and music score.
[0025] 2.2 Music score encoder selection: The music score encoder adopts a pre-trained MidiBERT-Piano model, which supports music score input in MIDI format;
[0026] 2.3 Feature Connector Selection: The feature connector adopts the Q-former structure to realize the alignment mapping between the musical score features extracted by the music score encoder and the audio features extracted by the large-scale audio language model.
[0027] In step 3, the specific steps for fine-tuning the evaluation model based on the large-scale audio language model SALMONN to support music score input can be as follows:
[0028] 3.1 Model Structure Confirmation: The structure of the Gezai Opera singing evaluation model comes from the large-scale audio language model SALMONN that supports musical score input in step 2. The model includes a dry audio encoder to be evaluated. The encoder consists of two parts: a Whisper encoder, which is responsible for encoding the content of the dry singing audio, and a BEATs encoder, which is responsible for encoding the acoustic feature information in the dry singing audio other than the content.
[0029] 3.2 Training Data and Objectives: The input to the evaluation model is the dry vocal audio of the samples from the dataset in Step 1 and the corresponding musical score. The output of the model is the comment text and the total evaluation score. The human comment text and the total evaluation score of the real samples are used as the standard answer for model training.
[0030] 3.3 Cross-entropy Loss Function Setting: For the evaluation text output by the singing performance evaluation model, the cross-entropy loss function is used for supervised training. The calculation formula is as follows:
[0031]
[0032] in, The cross-entropy loss function; For word morpheme number; The total number of words in the comment text; The first generation of the current evaluation model for Gezai opera singing style is generated by decoding. Each word element; To generate the first All words that the model has decoded and generated before the word element; To generate lexical units Under the condition of model generation, the first each word element The conditional probability; Logarithmic operations;
[0033] 3.4 The evaluation text output by the singing performance evaluation model finally includes a score. The word units in the score part are processed using the mean squared error loss function.
[0034]
[0035] Where m represents the total number of rating terms in the sample. To evaluate the scoring terms predicted by the model, These are the actual rating terms for the sample.
[0036] 3.5 Model Training: Using the training set in the dataset constructed in step 1.6 as the training data and the validation set as the validation data, the model training is performed with the sum of the cross-entropy loss and mean squared error loss functions defined in steps 3.3 and 3.4 as the optimization objective until the model converges.
[0037] In step 4, the specific steps for evaluating the singing style of Taiwanese opera applied to real-world scenarios and testing it on the test set are as follows:
[0038] 4.1 Input the dry audio and musical score of the test set samples into the vocal evaluation model trained in step 3. The output is the predicted comment text and score of the model.
[0039] 4.2 For the generation of comment text, the Bilingual Evaluation Understudy (BLEU) was used as the evaluation index; for the predicted scores, the Mean-Square Error (MSE), Linear Correlation Coefficient (LCC), Spearman Rank Correlation Coefficient (SRCC), and Kendall Tau Rank Correlation Coefficient (KTAU) were used as evaluation indexes.
[0040] Compared with the prior art, the outstanding advantages and technical effects of the present invention are as follows:
[0041] 1. In response to the problem that existing opera singing evaluation models can only provide MOS scores and lack interpretability, this invention changes the paradigm of Chinese opera singing evaluation technology. The proposed evaluation model can not only generate scores, but more importantly, it can generate comment text. The comment text generates textual descriptions of the strengths and weaknesses in terms of pitch, rhythm, range, timbre, singing techniques, volume variation, breath control, and emotional expression, so that users can understand the reasons for the score, which is conducive to further improvement of singing level.
[0042] 2. To address the lack of text-based evaluation datasets in the study of Chinese Gezai Opera singing evaluation, this invention constructs a Gezai Opera singing evaluation dataset with accompanying commentary texts. These commentary texts ensure that the evaluation process is well-founded and evidence-based. This dataset provides a data foundation for evaluating Gezai Opera singing based on large-scale audio language models.
[0043] 3. This invention innovatively adds a music score encoder to a large-scale audio language model, forming a dual-input model that can both "listen" to audio and "see" music scores, making the evaluation process well-founded and highly professional.
[0044] 4. In the evaluation technology of Chinese opera singing, this invention, when fine-tuning the evaluation model based on a large-scale audio language model, expands from the original cross-entropy loss to a combination of cross-entropy and mean squared error. This not only accurately generates singing evaluation text, but the mean squared error loss can also further constrain the accuracy of the score. Attached Figure Description
[0045] Figure 1 This invention constructs sample content corresponding to a single audio value in the dataset.
[0046] Figure 2 This is a diagram illustrating the overall architecture of the opera singing evaluation model proposed in this invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the following embodiments will be used in conjunction with the accompanying drawings to further illustrate the invention. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. For details not described in detail, conventional techniques in the art can be employed.
[0048] This invention first constructs a dataset for evaluating the singing style of Chinese Gezai Opera, containing textual comments. The dataset includes sentence-level dry vocal audio recordings, corresponding musical scores, evaluation text descriptions, and singing style evaluation scores. The audio evaluation descriptions come from professional and public reviewers, and cover eight aspects: pitch, rhythm, range, timbre, singing technique, volume variation, breath control, and emotional expression. After the dataset is constructed, the Gezai Opera singing style evaluation model is trained using a large-scale audio language model, SALMONN, with a new musical score encoder and low-rank adaptation (LoRA) fine-tuning. During model fine-tuning, the model is required to generate comments step-by-step according to the evaluation aspects; the text lexical units in this part use the cross-entropy loss function. Since the language model is required to output a total predicted score, a mean squared error loss function is added to the score lexical units. Finally, the trained model is applied to the evaluation of Chinese Gezai Opera singing styles, providing a unified and fair evaluation model for research on singing style synthesis in this field. This evaluation model changes the limitations of traditional methods that only provide scores, by generating text comments and understanding the reasons for the scores, which is conducive to further improving singing skills.
[0049] The specific implementation methods of this invention are as follows:
[0050] Step 1: Construct a dataset of Chinese Gezai Opera singing evaluations containing text comments;
[0051] 1.1 High-quality audio and corresponding scores were obtained from the Gezaixi (Chinese opera) dataset constructed by Zheng Meizhen et al. in "FT-GAN: Fine-Grained Tune Modeling for Chinese Opera Synthesis" (AAAI 2024, pp. 19697-19705) at the Thirty-Eighth Conference on Artificial Intelligence. Audio and scores were also obtained from the GOSMOS dataset for Gezaixi singing MOS prediction constructed by Bai Peng et al. in "Reference-free singing voice MOS prediction via multi-feature fusion, withintegrated feature analysis" (Vol. 241) in Applied Acoustics. Additionally, a number of audio recordings of ordinary students singing Gezaixi were recorded, and the corresponding scores were collected. The first part of the audio consists of 200 sentences of singing from professional Taiwanese opera actors in this embodiment; the second part consists of 500 sentences synthesized by the singing synthesis model in this embodiment; the third part consists of 300 sentences of singing from trainees in this embodiment; totaling 1000 sentence-level singing audios, saved in .wav file format, with a sampling rate of 16kHz, mono, ensuring the layering of singing levels and the diversity of audio sources. 1.2 For each sentence-level dry audio segment, a parallel sentence-level musical score file is matched. The musical score files are uniformly saved in MIDI format (.mid suffix) to ensure that the melody, rhythm and audio of the score are accurately aligned, with a sentence-level alignment accuracy of ≤50ms. A MusicXML format backup is also provided.
[0052] 1.3 Based on each vocal recording and its corresponding score, professional and public judges provided comments and scores for the audio recordings:
[0053] 1.3.1 The judging panel consisted of four professional Taiwanese opera judges with backgrounds in Taiwanese opera performance / teaching, and six Taiwanese opera enthusiasts as public judges.
[0054] 1.3.2 Reviewers will write comments for each audio segment based on eight dimensions: pitch, rhythm, vocal range, timbre, singing technique, volume variation, breath control, and emotional expression. A final score will be given, ranging from 1 to 5 points, with a 0.5-point interval. Each comment will be limited to a maximum of 200 Chinese characters.
[0055] The unified scoring criteria are as follows:
[0056] 1 point: The singing has serious pitch and rhythm errors, chaotic breathing, and cannot be properly understood. It does not have the basic charm of Taiwanese opera.
[0057] 2 points: There are many obvious pitch or rhythm errors in the singing, insufficient breath support, lack of emotional expression, and the overall singing effect is poor.
[0058] 3 points: The singing style is basically complete, with no obvious errors in pitch and rhythm, and a few acceptable flaws. It has the basic charm of Taiwanese opera.
[0059] 4 points: The singing is fluent and accurate, the pitch and rhythm are stable, the breath control is good, and the emotional expression is natural, with only very minor flaws.
[0060] 5 points: The singing is perfect and smooth, the pitch and rhythm are accurate, and the breath control, technique and emotional expression all reach the level of a professional actor, with no obvious flaws.
[0061] 1.3.3 After the evaluation process was completed, each audio segment corresponded to a musical score, 10 different reviewer comments, and 10 quantitative scores. The collected data structure is as follows: Figure 1 As shown.
[0062] 1.3.4 Optimization of Comment Text: GPT-4 was used to detect and polish the manually annotated comment text, maintaining the original meaning while expanding the diversity of text expression. For example, the original comment "Pitch is okay, rhythm is accurate, timbre is average" was polished to "Pitch basically matches the score, with occasional deviations at transitions; rhythm matches the style of Taiwanese opera, with no obvious rushing / dragging; timbre has basic charm, but the texture of the embellishment is insufficient."
[0063] 1.4 Ensure that the samples in the training set, validation set, and test set are not duplicated. Divide the dataset into subsets according to the ratio of 80% training set (800 sentences), 10% validation set (100 sentences), and 10% test set (100 sentences). This dataset is named "Chinese Gezai Opera Singing Evaluation Dataset with Text Comments".
[0064] Step 2: Construct a Gezai Opera singing evaluation model based on a large-scale audio language model;
[0065] 2.1 Based on the general large-scale audio language model SALMONN, an additional MIDI music score encoder and feature connector are added.
[0066] 2.2 The MIDI music score encoder uses a pre-trained MidiBERT-Piano model to accurately extract core features such as melody, rhythm, and pitch distribution from MIDI format music scores.
[0067] 2.3 A Q-former feature connector with two layers and 768 hidden dimensions is added after the MIDI music score encoder. This connector is used to connect and fuse the music score encoding features with the audio features of the large-scale audio language model SALMONN. The final structure of the Gezai Opera singing evaluation model is as follows: Figure 2 As shown.
[0068] Step 3: Fine-tune the evaluation model for Taiwanese opera singing style;
[0069] 3.1 The model input consists of sentence-level dry audio and corresponding musical scores from the data sample constructed in step 1. The model output consists of the predicted comment text and quantitative score.
[0070] 3.2 The required JSON data format for the model is:
[0071] "annotation": [
[0072] {
[0073] "wavpath": "1_0006.wav",
[0074] "midipath":"1_0006_score.mid",
[0075] The audio of this Taiwanese opera singing is generally accurate in pitch, but slightly lacking in key changes and sustained notes. The rhythm fits the melody, and the vocal range covers commonly used intervals, but the high and low notes lack expressiveness. The timbre has basic charm, and the singing technique mastered basic embellishment techniques, but the application was rigid. The volume has basic layering, but the contrast between loud and soft is not distinct. The breath support can support conventional singing, but the sustained notes are weak. The emotional expression fits the plot; overall score: 3.5.
[0076] "realscore": "3.5",
[0077] },
[0078] ...
[0079] ].
[0080] 3.3 The text instruction is as follows: You are an evaluator of Taiwanese opera singing style. Evaluate the singing audio in the following eight aspects in turn: pitch accuracy, rhythm, vocal range, timbre, singing technique, volume variation, breath control, and emotional expression. The text should end with a score. The total word count should not exceed 200 words.
[0081] 3.4 During model fine-tuning, the cross-entropy loss function is used to supervise the word generation process. The calculation formula is as follows:
[0082]
[0083] in, The first generation of the current evaluation model for Gezai opera singing style is generated by decoding. The model generated a total of [number] word units. Each word element, The cross-entropy loss function; For word morpheme number; The total number of words in the comment text; The first generation of the current evaluation model for Gezai opera singing style is generated by decoding. Each word element; To generate the first All words that the model has decoded and generated before the word element; To generate lexical units Under the condition of model generation, the first each word element The conditional probability; This is a logarithmic operation.
[0084] 3.5 During model fine-tuning, for the fractional word part, a mean squared error loss function is introduced to constrain prediction accuracy. The calculation formula is as follows:
[0085]
[0086] in, To evaluate the scoring terms predicted by the model, These are the actual rating terms for the sample.
[0087] 3.6 Low-rank adaptation (LoRA) fine-tuning of the model, only for the SALMONN decoder and Q-former feature connector, with the backbone encoder parameters frozen; parameter settings: lora_rank is 8, lora_alpha is 32, and lora_dropout is 0.1.
[0088] 3.7 The core configurations of the optimizer and learning rate scheduler in the experiment are as follows: the maximum number of training epochs is set to 10; the learning rate warm-up steps are 3000 steps, and the initial warm-up learning rate is 1×10. -6 The initial learning rate is 3×10 -5 The minimum learning rate is 1×10. -5 The weight decay factor was set to 0.05, and the optimizer's beta2 parameter was set to 0.999. The model was trained using a single Nvidia A40 graphics card.
[0089] 3.8 Training convergence condition: Training is stopped when the validation set BLEU-4≥45, MSE≤0.3, and LCC≥0.85.
[0090] Step 4: Evaluation and testing of the singing style of Taiwanese opera audio;
[0091] 4.1 Input the test set divided in step 1 into the fine-tuned and trained Gezai Opera singing evaluation model in step 3. The model outputs an aspect-level evaluation text containing 8 dimensions and an overall evaluation text paragraph with a total score.
[0092] 4.2 For the comment text, calculate the BLEU (Bilingual Assessment Auxiliary) index to evaluate the consistency between the generated text and the manually annotated text; for the rating, calculate the MSE (Mean Squared Error), LCC (Linear Correlation Coefficient), SRCC (Spearman Rank Correlation Coefficient), and KTAU (Kendall Rank Correlation Coefficient) indices to evaluate the correlation between the predicted rating and the actual rating, and test the performance of the model.
[0093] Indicator Results:
[0094] Text: BLEU-4=48.2, ROUGE-L=71.5, which is highly consistent with human evaluations;
[0095] Scoring: MSE=0.28, LCC=0.88, SRCC=0.86, KTAU=0.79, showing a strong correlation with human scoring;
[0096] Comparative experiment: Without sheet music input, LCC dropped to 0.72 and BLEU-4 dropped to 39.5, demonstrating the effect of bimodal sheet music input on improving evaluation accuracy.
[0097] The main innovations of this invention are as follows: First, it constructs the first Chinese Gezai Opera singing evaluation dataset containing multi-dimensional comment texts, filling a data gap in this field; second, it innovatively combines the large-scale audio language model SALMONN with the MidiBERT-Piano music score encoder and Q-former feature connector to achieve dual-modal evaluation by listening to audio and reading music score; third, it proposes a LoRA fine-tuning strategy using a joint loss of cross-entropy and mean squared error, constraining the model to generate comments and total scores step by step according to eight dimensions, ensuring both the professionalism and coherence of the comment texts and improving the accuracy and interpretability of the scores; fourth, it breaks through the limitation of traditional opera evaluations that only provide a single MOS score, realizing a paradigm shift in evaluation by combining scoring and interpretable comments, providing a unified, fair, and reproducible automated evaluation tool for Gezai Opera singing teaching and synthetic model evaluation.
[0098] All equivalent changes and improvements made within the scope of this invention application shall still fall within the patent coverage of this invention.
Claims
1. A method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model, characterized in that... Includes the following steps: Step 1: Construct a Chinese Gezai Opera singing evaluation dataset containing text comments. The dataset includes sentence-level dry vocal audio, the corresponding musical score, the evaluation text description, and the total singing evaluation score. Step 2: Based on the large-scale audio language model SALMONN, a music score encoder and feature connector are added to build a Gezai Opera singing evaluation model that can receive dual-modal input of dry vocal audio and music score. Step 3: Low-rank adaptive fine-tuning is adopted. During the model fine-tuning process, the model is constrained to generate the evaluation text step by step according to multiple singing evaluation dimensions. The evaluation text word units adopt the cross-entropy loss function. Since the language model is required to output the total evaluation prediction score, the mean squared error loss function is added to the word unit part of the evaluation score. Step 4: Apply the trained Gezai Opera singing evaluation model to the automatic evaluation of Gezai Opera singing audio, and output multi-dimensional evaluation text and total score.
2. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 1, the sentence-level dry vocal audio sources include audio from professional actors, audio from amateur singers, and synthesized audio generated by a Gezai opera vocal synthesis model. The recording devices include professional microphones or mobile phones.
3. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 1, the musical score corresponding to the audio is a standard musical score collected for each sentence-level dry audio, in MIDI or MusicXML format, with the score aligned with the audio sentence level.
4. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 1, the evaluation text description corresponding to the audio is collected. The evaluation description text of each audio file is marked by professional reviewers and public reviewers. The evaluation dimensions are multiple preset singing evaluation dimensions, including eight aspects: pitch, rhythm, vocal range, timbre, singing technique, volume variation, breath control, and emotional expression.
5. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 1, the total score for the singing evaluation is determined by collecting each audio file and manually annotating it using a standardized crowdsourcing evaluation process and unified scoring standards. The manually annotated comments are then detected and refined using a large text model to expand the diversity of the comments while maintaining the original meaning. The Chinese Gezai Opera singing evaluation dataset is then divided into a training set, a validation set, and a test set.
6. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 2, after the large-scale audio language model SALMONN adds a music score encoder, the input includes the audio to be evaluated, the corresponding music score, and the evaluation instructions. The output includes aspect-level comments and a total score.
7. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 2, the music score encoder uses a pre-trained MidiBERT-Piano model, and the feature connector uses a Q-former structure to fuse and align the music score encoding features with the audio features of the large-scale audio language model SALMONN.
8. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 2, the input of the Gezai Opera singing evaluation model includes the dry vocal audio to be evaluated, the corresponding standard musical score, and the evaluation instructions. The output includes sequential evaluation text for multiple singing evaluation dimensions and a total score.
9. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 3, the generation of the described text terms uses the cross-entropy loss function, which is calculated using the following formula: in, The cross-entropy loss function; For word morpheme number; The total number of words in the comment text; The first generation of the current evaluation model for Gezai opera singing style is generated by decoding. Each word element; To generate the first All words that the model has decoded and generated before the word element; To generate lexical units Under the condition of model generation, the first each word element The conditional probability; This is a logarithmic operation.
10. The method for evaluating the singing style of Chinese Gezai Opera based on a large-scale audio language model according to claim 1, characterized in that... In step 3, for the fractional word part, a mean squared error loss function constraint is added. The formula for calculating the mean squared error loss function is as follows: in, Let the mean squared error loss function be used. Indicates the total number of rating terms. For word morpheme number; To evaluate the scoring terms predicted by the model, These are the actual rating terms for the sample.