AI-driven playing quality evaluation generation method and device

Through the AI-driven multi-dimensional evaluation method, combined with feature extraction and visual feedback, the problem of inaccurate evaluation in the existing technology is solved, more accurate performance quality evaluation and personalized guidance are achieved, and users' practice effects are improved.

CN120299479APending Publication Date: 2025-07-11LEHE DATA INFORMATION TECH JIANGSU CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510526620.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

现有技术在演奏评估中缺乏多维度分析,导致评分结果与实际演奏水平不符,影响用户对演奏表现的准确理解与有效提升。

Method used

Using an AI-driven multi-dimensional fusion evaluation method, through feature extraction and visual feedback, combined with user-selected repertoire and instrument types, personalized performance quality evaluation and guidance are provided.

Benefits of technology

It improves the accuracy of evaluation, enhances the pertinence of practice guidance, and improves the efficiency of users' performance level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299479A_ABST
    Figure CN120299479A_ABST
Patent Text Reader

Abstract

The invention provides an AI-driven playing quality evaluation generation method and device, and relates to the technical field of playing quality evaluation, and the method comprises the steps: providing a track for user selection, and collecting a playing audio based on the track selected by a user; performing feature extraction on the playing audio to obtain playing audio features; performing quality evaluation on the playing audio characteristics to obtain an audio quality evaluation result; and performing visual feedback according to an audio quality evaluation result. According to the method and the device, the technical problems that the scoring result is inconsistent with the actual playing level and the accurate understanding and effective improvement of the user on the playing performance are further influenced due to the fact that the evaluation means is single and static and the analysis of multi-dimensional elements such as playing emotion and style is lacked in the prior art can be solved; the technical target of intelligent playing quality judgment based on combination of multi-dimensional fusion evaluation and personalized feedback is achieved, and the technical effects of improving evaluation accuracy, enhancing user practice guidance pertinence and improving playing level progress efficiency are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of performance quality evaluation, and particularly to a method and device for generating performance quality evaluation driven by AI. Background Art

[0002] With the rapid development of technologies such as artificial intelligence, speech recognition, and music information retrieval, more and more music education assistance systems are widely used in daily learning and training scenarios, especially outstanding in aspects such as instrument performance grading simulation, skill scoring, and intelligent feedback. Common music grading applications on the current market already support functions such as performance recording, scoring evaluation, and reference demonstration for various instruments including piano, guzheng, and violin, providing a convenient auxiliary learning platform for music beginners and professional candidates. However, there are still many deficiencies in the existing technologies in terms of evaluation accuracy, degree of feedback intelligence, and personalized recommendation.

[0003] Currently, in terms of existing performance evaluation dimensions, most systems rely on traditional static algorithms for pitch and rhythm matching, lacking dynamic and multi-dimensional comprehensive analysis means. Most applications only stay at the basic note matching level, and it is difficult to quantitatively evaluate deep-level performances such as the emotional expression, performance style, and dynamics changes of performers, resulting in a deviation between the scoring results and the real performance quality. Secondly, in terms of feedback content, most systems only give simple score feedback, lacking graphical analysis and bar-by-bar problem positioning functions, making it difficult for users to intuitively understand the performance problems, reducing the improvement efficiency and practice pertinence. In addition, in terms of repertoire recommendation, existing platforms usually display the content of the repertoire library in a fixed grading manner, failing to perform personalized matching in combination with the user's performance purpose, instrument type, technical level, etc., resulting in a large information screening burden for users when selecting practice repertoires.

[0004] In summary, there are technical problems in the existing technologies that due to the single and static evaluation means and the lack of analysis of multi-dimensional elements such as performance emotion and style, the scoring results do not match the actual performance level, further affecting the user's accurate understanding and effective improvement of performance. Summary of the Invention

[0005] The purpose of this application is to provide a method and device for generating performance quality evaluation driven by AI to solve the technical problems in the existing technologies that due to the single and static evaluation means and the lack of analysis of multi-dimensional elements such as performance emotion and style, the scoring results do not match the actual performance level, further affecting the user's accurate understanding and effective improvement of performance.

[0006] In view of the above problems, this application provides a method and device for generating performance quality evaluation driven by AI.

[0007] In a first aspect, the present application provides a method for generating a performance quality evaluation under driving, which is implemented by a device for generating a performance quality evaluation under driving, and includes: providing pieces for user selection, and collecting performance audio based on the user-selected piece; extracting features from the performance audio to obtain performance audio features; evaluating the quality of the performance audio features to obtain an audio quality evaluation result; and performing visual feedback according to the audio quality evaluation result.

[0008] In a second aspect, the present application further provides a device for generating a performance quality evaluation under driving, which is used to execute the method for generating a performance quality evaluation under driving as described in the first aspect, and includes: an audio acquisition module for providing pieces for user selection and collecting performance audio based on the user-selected piece; a feature obtaining module for extracting features from the performance audio to obtain performance audio features; a result obtaining module for evaluating the quality of the performance audio features to obtain an audio quality evaluation result; and a visual feedback module for performing visual feedback according to the audio quality evaluation result.

[0009] The technical solution provided in the present application has at least the following technical effects or advantages: By achieving the technical goal of intelligent performance quality determination combining multi-dimensional fusion evaluation and personalized feedback, the technical effects of improving evaluation accuracy, enhancing the pertinence of user practice guidance, and improving the efficiency of performance level progress are achieved.

[0010] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives the specific implementation manners of the present application. It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only exemplary, and for those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0012] Figure 1 It is a schematic flowchart of the method for generating a performance quality evaluation under driving of the present application;

[0013] Figure 2 It is a schematic structural diagram of the device for generating a performance quality evaluation under driving of the present application.

[0014] Description of the attached drawing reference numerals: Audio acquisition module 11, feature acquisition module 12, result obtaining module 13, visual feedback module 14. Detailed implementation manners

[0015] By providing a method and device for generating a performance quality evaluation driven by AI, the present application solves the technical problem in the prior art that due to the single and static evaluation means and the lack of analysis of multi-dimensional elements such as performance emotion and style, the scoring result does not match the actual performance level, further affecting the user's accurate understanding and effective improvement of the performance. The technical goal of realizing intelligent performance quality determination by combining multi-dimensional fusion evaluation and personalized feedback is achieved, and the technical effects of improving the evaluation accuracy, enhancing the pertinence of user practice guidance, and improving the efficiency of performance level progress are achieved.

[0016] Next, the technical solutions in the present application will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application. In addition, it should be noted that for the convenience of description, only the parts related to the present application are shown in the accompanying drawings rather than all.

[0017] Embodiment 1. Please refer to the attached Figure 1 , the present application provides a method for generating a performance quality evaluation driven by [driver], which is applied to a device for generating a performance quality evaluation driven by [driver], and specifically includes the following steps:

[0018] S1: Provide a repertoire for the user to select, and collect the performance audio based on the user-selected repertoire.

[0019] Specifically, a list of repertoires for selection is presented to the user, including a variety of repertoires covering different difficulty levels, styles or performance requirements, for the user to select according to personal preferences and needs. The user can choose the repertoire they are interested in for performance and practice.

[0020] Then, according to the user's selection in the repertoire list, the audio content of the user's performance is recorded and collected through a microphone or other audio acquisition devices, etc., to capture the audio signal during the user's performance.

[0021] S2: Extract features from the performance audio to obtain performance audio features.

[0022] Specifically, analyzing the audio played by the user and extracting data representing the audio features, which can reflect the key information in the audio, such as pitch, volume, rhythm, sound quality, etc. Feature extraction can transform complex audio signals into digital and easily analyzable feature information, obtaining the features of the played audio, including pitch time series, spectrogram, rhythm information, etc., enabling a comprehensive understanding of the performance in aspects such as intonation, rhythm, and emotion of the performance.

[0023] S3: Perform a quality assessment on the features of the played audio to obtain an audio quality assessment result.

[0024] Specifically, evaluate the quality of the performance based on the audio features, and judge whether the accuracy, fluency, and emotional expression of the performance meet the expected standards. The quality assessment measures the performance of each feature according to preset criteria, and then combines data from multiple dimensions to obtain a comprehensive evaluation of the audio quality, reflecting the quality of the played audio.

[0025] S4: Provide visual feedback according to the audio quality assessment result.

[0026] Specifically, according to the result of the audio quality assessment, present the evaluated content to the user in a graphical and textual manner. The audio quality assessment result usually includes scores in aspects such as intonation, rhythm, and emotional expression, and the visual feedback displays these score results in the form of charts, graphs, etc., to help the user more intuitively understand their own performance.

[0027] Furthermore, this application also includes: receiving performance purpose information and outputting a matching list of pieces; receiving the piece selected by the user from the list of pieces as the user-selected piece; outputting the audio to be played according to the user-selected piece; and performing audio acquisition based on the audio to be played to obtain the played audio.

[0028] Specifically, obtain the purpose for which the user wants to perform or practice, such as for exam simulation, daily practice, or skill enhancement, etc., and then provide more accurate content matching according to different learning needs.

[0029] Then, according to the performance purpose provided by the user, automatically screen and generate a collection of relevant pieces for the user to choose from, which may be screened according to multiple dimensions such as the exam syllabus, difficulty level, and skill type. For example, after the user sets "Grade 8 of piano" as the goal, pieces such as Liszt's Etudes and Chopin's Waltzes may be provided for the user to select works from.

[0030] After that, the user selects one or more pieces of music from the recommended options as the target pieces for this performance, and then analyzes and grades them according to the score examples and performance specifications of the selected pieces. For example, if the user selects the guzheng solo version of "Butterfly Lovers" from the recommended list, all subsequent pitch and rhythm analyses will refer to the standard performance of this piece.

[0031] Subsequently, the standard performance version of the selected piece (such as MIDI or an expert performance recording) is played to the user as a reference audio for practice imitation or comparison, which includes standard rhythm, pitch, and emotional expression.

[0032] Finally, the user's performance is recorded in real time through a microphone or other acquisition device to generate the user's performance audio. For example, the user uses the mobile phone microphone to record a 3-minute guzheng performance as the performance audio. Table 1 shows the processing instructions for collecting performance audio based on the user's selected piece.

[0033] Table 1: Processing Instructions for Collecting Performance Audio Based on the User's Selected Piece

[0034]

[0035] Furthermore, this application also includes: equipping each piece in the piece list with an expert performance as the gold standard; obtaining the corresponding gold standard according to the user's selected piece to generate the target gold standard; splitting the notes of the target gold standard to obtain the starting points of standard notes; converting the starting points of the standard notes into spectrograms to obtain standard spectrograms; extracting the pitch of the target gold standard in a time series to obtain a standard pitch time series; and combining the standard spectrograms and the standard pitch time series to obtain standard audio features.

[0036] Specifically, prepare a reference performance version for each recommended piece. The reference version can be the version of the piece performed by experienced performers and post-processed to ensure that the pitch, rhythm, and emotional expression are all in an optimal state, and is used as a reference for comparing the user's performance quality. Among them, the reference performance version can be customized by professionals in this field.

[0037] Next, when the user selects a certain piece from the piece list, an expert performance audio corresponding to it is automatically retrieved to ensure that there is an accurate and consistent reference object for the user's performance analysis.

[0038] Then, the audio of the expert performance is analyzed through audio processing technologies such as mutation detection or Onset detection technology to identify the starting position of each note on the time axis, which is used as note splitting to capture the hitting moment of each note.

[0039] Subsequently, using the starting point of the extracted notes as an anchor, the entire audio is converted into a spectrogram, that is, a two-dimensional image transformed from the time-waveform domain to the time-frequency domain, showing the frequency components contained at each moment, which helps analyze the pitch, resonance, and intensity changes of the notes. For example, in a 3-minute piece of music, a spectrogram containing 3000 frames may be generated, with each frame corresponding to a note starting point.

[0040] Next, through pitch recognition algorithms such as YIN or CREPE, the dominant frequency at each moment, that is, the pitch value, is continuously extracted from the audio and arranged in chronological order to form a continuous sequence, enabling the determination of whether a certain performance is in an ascending, descending, or stable state.

[0041] Finally, the standard spectrogram and the standard pitch time series are combined to obtain the standard audio features, which are fused into a complete feature set for subsequent comparison and analysis with the user's performance.

[0042] Furthermore, this application also includes: receiving performance instrument information, performing piece matching in combination with the performance purpose information, outputting a first piece list, and adding the first piece list to the piece list, where the first piece list includes first standard audio features based on the performance instrument information.

[0043] Specifically, obtain the type of instrument currently used by the user, such as a guzheng, piano, violin, etc. Different instruments have different performance techniques, timbre characteristics, and corresponding piece libraries, and targeted content needs to be provided according to the differences in instrument types.

[0044] Next, on the basis of identifying the instrument type, further consider the user-set practice or grading goal, so as to more accurately screen out suitable pieces. For example, when it is known that the user is using a piano and the goal is to pass the seventh grade, give priority to matching piano works that meet the seventh-grade requirements, such as Beethoven's Sonatina or Czerny's Etudes, rather than recommending guzheng pieces or beginner-level pieces.

[0045] Subsequently, display the matched pieces in the form of a list for the user to select or refer to. The first piece list is the result of double screening by instrument type and performance purpose, and has stronger pertinence and practicality.

[0046] Next, add the first piece list to the piece list, and integrate the current recommendation result with the user's original piece library to form an updated complete piece set.

[0047] Among them, the first piece list includes first standard audio features based on the performance instrument information, accompanied by standardized audio features related to this instrument, including spectrograms of expert performance audio, pitch time series, etc., which can provide standard examples for users to imitate and compare.

[0048] Furthermore, this application also includes: performing pitch accuracy evaluation on the performance audio features based on the standard audio features to obtain a pitch accuracy evaluation coefficient; performing rhythm evaluation on the performance audio features according to the standard audio features to obtain a rhythm evaluation coefficient; performing emotional expression evaluation through the standard audio features and the performance audio features to generate an emotional expression evaluation coefficient; and performing comprehensive performance quality evaluation based on the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient to obtain the audio quality evaluation result.

[0049] Specifically, compare the pitch change data of the user's performance with the pitch sequence of the expert's standard performance, and calculate the deviation degree between each note in the user's performance and the corresponding standard pitch. The pitch accuracy evaluation coefficient is a value reflecting the overall pitch accuracy. The closer the value is to the ideal value, the more accurate the pitch.

[0050] Then, compare the start time of each note in the user's performance with the start point of the note in the standard performance to measure the stability and accuracy of the rhythm. The rhythm evaluation coefficient is an index measuring the note alignment degree and the time value control ability. The closer the value is to the ideal value, the more accurate the rhythm. For example, if the user frequently starts early or late in a certain piece of music, and each note deviates from the standard rhythm point by more than 200 milliseconds, the rhythm score may be evaluated as 0.6.

[0051] Next, use dimensions such as volume dynamic change, speed fluctuation, and timbre tension to compare the consistency of emotional fluctuations between the user's performance and the expert's performance. Emotional expression not only includes the contrast of note strength, but also includes the coherence of emotional transitions in the whole performance. Calculate a coefficient of emotional matching degree by analyzing signal indicators such as volume envelope and beat stretching.

[0052] Finally, perform comprehensive performance quality evaluation based on the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient, that is, weighted fusion to obtain the audio quality evaluation result, and form a score of the overall performance. The weighting method can be adjusted according to the performance level or user requirements.

[0053] Furthermore, this application also includes: configuring the weights of the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient; weighting the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient according to the weights of the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient to obtain the audio quality evaluation result.

[0054] Specifically, weights are set for the importance of the three evaluation dimensions, namely, the weight of the pitch evaluation coefficient, the weight of the rhythm evaluation coefficient, and the weight of the emotional expression evaluation coefficient, which are used for weighted calculation to reflect the proportion of each index in the audio quality evaluation result. For example, at the beginner stage, the pitch may be assigned a weight of 0.5, while the rhythm and emotion are 0.3 and 0.2 respectively; at the advanced stage, the weight of the emotional expression may be increased to 0.4. The configuration of different weights can be adjusted according to different learning stages, instrument types, or user personalized goals, so as to achieve a more flexible scoring system that conforms to the logic of music education.

[0055] According to the weight of the pitch evaluation coefficient, the weight of the rhythm evaluation coefficient, and the weight of the emotional expression evaluation coefficient, the pitch evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient are weighted to obtain the audio quality evaluation result. Multiply the actual score of each scoring dimension by the corresponding weight, and then add the three weighted scores to get a final comprehensive score. For example, in the performance of a piano piece by a certain user, the pitch score is 0.85, the rhythm is 0.75, and the emotion is 0.9; if the configured weights are 0.4, 0.3, and 0.3 respectively, the comprehensive score is 0.85×0.4 + 0.75×0.3 + 0.9×0.3, and the calculation result is approximately 0.83. Converting this to a percentage system is 83 points, and it may be marked as "mid-advanced level".

[0056] Furthermore, the present application further includes: presenting a multi-dimensional score based on the audio quality evaluation result to obtain a performance score; extracting the performance purpose information and determining whether there is a mock exam purpose in the performance purpose; if so, converting the performance score to an exam level to obtain a performance level; combining the selected score sheet in the user-selected piece, and generating highlighted incorrect notes, error time annotation of the rhythm diagram, and emotional fluctuation curve based on the audio quality evaluation result to obtain visual feedback information; and performing user interface feedback based on the performance level and the visual feedback information.

[0057] Specifically, based on the weighted results of the evaluation coefficients such as pitch, rhythm, and emotion obtained, a multi-dimensional comprehensive score is performed. When viewing the score, the user can not only see an overall score but also clearly understand the performance of each score.

[0058] Then, it is judged according to the user's preset goals (such as exam, practice, performance, etc.). If the user's goal is a mock exam, it automatically enters the corresponding evaluation mode and adjusts the criteria for scoring and feedback. For example, if the user's purpose is to take a grade examination, strict scoring and suggestions are made according to the standards and scoring requirements of the grade examination.

[0059] After determining that the user's goal is a mock exam, convert the performance score into the corresponding exam level according to the evaluation result. For example, if the audio quality evaluation result is 85 points, when taking the mock level 9 exam, it may be converted to "qualified for level 8" or "not up to the standard for level 9", so that the user can clearly understand the gap between their current level and the goal.

[0060] Then, according to the sheet music content selected by the user, compare the differences between the performance and the standard performance, and display the note errors, rhythm deviations, and deficiencies in emotional expression in a graphical way. For example, the incorrect notes in the performance can be highlighted on the sheet music, the early or late time in rhythm can be marked, and the volume change can be shown through a curve graph to help the user more intuitively understand the problems existing in the performance.

[0061] Finally, combine the performance score and the visual feedback and provide the user with intuitive and clear interface feedback. For example, after the user completes the performance, the specific score and improvement suggestions will be displayed on the interface, and the information such as the positions of the incorrect notes and the rhythm errors may also be presented through colors, charts, etc., to help the user understand their strengths and weaknesses and thus formulate a more effective practice plan.

[0062] Furthermore, this application also includes: if not, perform user interface feedback based on the performance score and the visual feedback information.

[0063] Specifically, if the user does not set the goal of a mock exam, use the performance score and the visual feedback information for the display of the user interface. The performance score refers to the score evaluated according to dimensions such as intonation, rhythm, and emotion, reflecting the overall level of the user's performance. The visual feedback information includes the highlighting of incorrect notes, the marking of the rhythm error time, the emotional fluctuation curve, etc., which are presented in a graphical and colorized way to help the user intuitively understand their performance. For example, if the intonation score of the user's performance is 0.85, the rhythm score is 0.75, and the emotion score is 0.8, a comprehensive score of 0.8 will be displayed, and by highlighting the incorrect notes in the performance, marking the rhythm deviation time, and drawing the emotional change curve, etc., to help the user accurately identify the areas that need improvement in the performance. In addition, the interface can also provide targeted improvement suggestions, such as "The pitch of note B is slightly off in the 12th measure, it is recommended to practice pitch control more" or "The rhythm is slightly behind in the 20th measure, pay attention to aligning with the metronome".

[0064] Furthermore, this application also includes: suppressing noise in the performance audio to obtain noise-reduced performance audio; trimming the first blank part and the second blank part in the noise-reduced performance audio to obtain silence-trimmed performance audio; normalizing the audio of the silence-trimmed performance audio to obtain standard-volume performance audio; segmenting the notes in the standard-volume performance audio to obtain the starting points of performance notes; converting the starting points of the performance notes into a spectrogram to obtain a performance spectrogram; extracting the pitch of the standard-volume performance audio in a time series to obtain a performance pitch time series; and combining the performance spectrogram and the performance pitch time series to obtain the performance audio features.

[0065] Specifically, unnecessary background noise in the audio, such as ambient sound, wind sound, or noise, is removed through filtering to improve the clarity of the audio, facilitating subsequent analysis. For example, if the user performs in a noisy environment, noise suppression can effectively reduce background noise, thereby more accurately extracting the note information of the performance.

[0066] Next, in the noise-reduced audio, the blank parts before the start and after the end of the performance, as well as the pauses during the performance, are removed. By trimming the blank parts and pauses, unnecessary data during analysis can be reduced, making the audio more compact and focusing on the actual performance content.

[0067] Then, the volume of the audio is adjusted to a standard range. Audio normalization technology can make the maximum volume of the audio reach a preset value, ensuring that the volume of the audio is neither too low nor too high when playing. For example, if the volume of some parts is low while that of other parts is high, normalization processing can balance the difference, making the volume of the entire audio more uniform.

[0068] Next, the start time point of each note is identified through audio analysis. Note segmentation is used to accurately locate the start position of each note, and the time information of each note in the performance can be accurately judged.

[0069] Then, through frequency analysis of the note starting points and conversion into a spectrogram, the audio signal is visually displayed in the frequency domain, showing the intensity of the audio at different frequencies. The spectrogram can help understand the frequency components of the notes, identify the pitch and timbre.

[0070] Next, by analyzing the time and frequency information in the audio, the pitch changes of the notes are extracted, the pitch corresponding to each note in the performance is obtained, and a time series is generated to show the changes of each note during the performance.

[0071] Finally, the performance audio features are obtained by combining the performance spectrogram and the performance pitch time series, forming a complete audio feature representation, reflecting information such as the timbre, intonation, and rhythm of the audio, and providing a basis for subsequent audio analysis and scoring.

[0072] In summary, the method for generating a performance quality assessment under driving provided by the present application has the following technical effects: By achieving the technical goal of intelligent performance quality determination that combines multi-dimensional fusion assessment and personalized feedback, the technical effects of improving the assessment accuracy, enhancing the pertinence of user practice guidance, and improving the efficiency of performance level improvement are achieved.

[0073] Embodiment 2, based on the same inventive concept as the method for generating a performance quality assessment under driving in the foregoing embodiment, the present application further provides a device for generating a performance quality assessment under driving. Please refer to the attached Figure 2 , including: an audio acquisition module, configured to provide pieces for user selection and acquire performance audio based on the user-selected piece; a feature obtaining module, configured to extract features from the performance audio to obtain performance audio features; a result obtaining module, configured to perform quality assessment on the performance audio features to obtain an audio quality assessment result; and a visualization feedback module, configured to perform visualization feedback according to the audio quality assessment result.

[0074] Furthermore, the device for generating a performance quality assessment under driving is further configured to: receive performance purpose information and output a matching piece list; receive the piece selected by the user from the piece list as the user-selected piece; output the to-be-played audio according to the user-selected piece; and perform audio acquisition based on the to-be-played audio to obtain the performance audio.

[0075] Furthermore, the device for generating a performance quality assessment under driving is further configured to: equip each piece in the piece list with an expert performance as the gold standard; obtain the corresponding gold standard according to the user-selected piece to generate a target gold standard; perform note segmentation on the target gold standard to obtain the starting point of standard notes; convert the starting point of standard notes into a spectrogram to obtain a standard spectrogram; extract the pitch of the target gold standard in a time series to obtain a standard pitch time series; and combine the standard spectrogram and the standard pitch time series to obtain standard audio features.

[0076] Furthermore, the device for generating a performance quality assessment under driving is further configured to: receive performance instrument information, perform piece matching in combination with the performance purpose information, output a first piece list, and add the first piece list to the piece list, where the first piece list includes first standard audio features based on the performance instrument information.

[0077] Furthermore, the performance quality evaluation generation device under the drive is further configured to: perform pitch evaluation on the performance audio features based on the standard audio features to obtain a pitch evaluation coefficient; perform rhythm evaluation on the performance audio features according to the standard audio features to obtain a rhythm evaluation coefficient; perform emotional expression evaluation through the standard audio features and the performance audio features to generate an emotional expression evaluation coefficient; perform comprehensive performance quality evaluation according to the pitch evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient to obtain the audio quality evaluation result.

[0078] Furthermore, the performance quality evaluation generation device under the drive is further configured to: configure weights for the pitch evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient; weight the pitch evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient according to the weights of the pitch evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient to obtain the audio quality evaluation result.

[0079] Furthermore, the performance quality evaluation generation device under the drive is further configured to: perform multi-dimensional score presentation according to the audio quality evaluation result to obtain a performance score; extract the performance purpose information, and determine whether there is a mock exam purpose in the performance purpose; if so, perform exam grade conversion on the performance score to obtain a performance grade; combine the score sheet selected from the user-selected repertoire, and perform highlighting of wrong notes, error time annotation of the rhythm diagram, and generation of an emotional fluctuation curve according to the audio quality evaluation result to obtain visual feedback information; perform user interface feedback based on the performance grade and the visual feedback information.

[0080] Furthermore, the performance quality evaluation generation device under the drive is further configured to: if not, perform user interface feedback based on the performance score and the visual feedback information.

[0081] Furthermore, the performance quality evaluation generation device under the drive is further configured to: perform noise suppression on the performance audio to obtain a noise-reduced performance audio; perform trimming on the first blank part and the second blank part in the noise-reduced performance audio to obtain a silent-trimmed performance audio; perform audio normalization on the silent-trimmed performance audio to obtain a standard-volume performance audio; perform note segmentation on the standard-volume performance audio to obtain the starting points of performance notes; perform spectrogram conversion on the starting points of performance notes to obtain a performance spectrogram; perform pitch extraction of the time series on the standard-volume performance audio to obtain a performance pitch time series; combine the performance spectrogram and the performance pitch time series to obtain the performance audio features.

[0082] In the present specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The method and specific examples for generating a performance quality assessment under driving in the foregoing first embodiment are equally applicable to the apparatus for generating a performance quality assessment under driving in this embodiment. Through the foregoing detailed description of the method for generating a performance quality assessment under driving, those skilled in the art can clearly understand the apparatus for generating a performance quality assessment under driving in this embodiment. Therefore, for the sake of brevity of the specification, no further details will be provided here.

[0083] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0084] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A method and device for generating an evaluation of performance quality driven by AI, characterized in that, including: Providing tracks for user selection and collecting performance audio based on the user-selected tracks; Extracting features from the performance audio to obtain performance audio features; Evaluating the quality of the performance audio features to obtain an audio quality evaluation result; Performing visual feedback according to the audio quality evaluation result.

2. The AI-driven performance quality evaluation generation method and device according to claim 1, wherein Providing tracks for user selection and collecting performance audio based on the user-selected tracks, including: Receiving performance purpose information and outputting a matching track list; Receiving the selected track by the user in the track list as the user-selected track; Outputting the audio to be performed according to the user-selected track; Performing audio collection based on the audio to be performed to obtain the performance audio.

3. The AI-driven performance quality evaluation generation method and device according to claim 2, wherein Before evaluating the quality of the performance audio features to obtain an audio quality evaluation result, including: Equipping each track in the track list with an expert performance as the gold standard; Obtaining the corresponding gold standard according to the user-selected track to generate a target gold standard; Segmenting the notes of the target gold standard to obtain the starting points of standard notes; Converting the starting points of the standard notes into spectrograms to obtain standard spectrograms; Extracting the pitch of the target gold standard in time series to obtain a standard pitch time series; Combining the standard spectrogram and the standard pitch time series to obtain standard audio features.

4. The AI-driven performance quality evaluation generation method and device according to claim 3, wherein Before evaluating the quality of the performance audio features to obtain an audio quality evaluation result, further including: Receiving performance instrument information, performing track matching in combination with the performance purpose information, outputting a first track list, and adding the first track list to the track list, where the first track list includes first standard audio features based on the performance instrument information.

5. The AI-driven performance quality evaluation generation method and device according to claim 3, characterized in that Evaluating the quality of the performance audio features to obtain an audio quality evaluation result, including: Performing pitch accuracy evaluation on the performance audio features based on the standard audio features to obtain a pitch accuracy evaluation coefficient; Performing rhythm evaluation on the performance audio features according to the standard audio features to obtain a rhythm evaluation coefficient; Performing emotional expression evaluation through the standard audio features and the performance audio features to generate an emotional expression evaluation coefficient; Performing comprehensive performance quality evaluation according to the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient to obtain the audio quality evaluation result.

6. The method and device for generating a performance quality evaluation driven by AI according to claim 5, characterized in that, Performing comprehensive performance quality evaluation according to the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient to obtain the audio quality evaluation result, including: Configuring weights for the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient; Weighting the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient according to the weights of the pitch accuracy evaluation coefficient, the rhythm evaluation coefficient, and the emotional expression evaluation coefficient to obtain the audio quality evaluation result.

7. The method and device for generating a performance quality evaluation driven by AI according to claim 2, characterized in that, Performing visual feedback according to the audio quality evaluation result, including: Presenting multi-dimensional scores according to the audio quality evaluation result to obtain a performance score; Extracting the performance purpose information and judging whether there is a simulated exam purpose in the performance purpose; If it exists, convert the performance score to an examination level to obtain a performance level. Combine the selected score surface from the user-selected repertoire, and based on the audio quality evaluation result, highlight incorrect notes, mark the error time of the rhythm diagram, and generate an emotional fluctuation curve to obtain visual feedback information. Perform user interface feedback based on the performance level and the visual feedback information.

8. The AI-driven performance quality evaluation generation method and device according to claim 7, wherein Performing visual feedback according to the audio quality evaluation result further includes: If it does not exist, perform user interface feedback based on the performance score and the visual feedback information.

9. The method for generating an evaluation of performance quality driven by AI according to claim 1, wherein, Extract features from the performance audio to obtain performance audio features, including: Suppress noise in the performance audio to obtain a noise-reduced performance audio. Clip the first blank part and the second blank part in the noise-reduced performance audio to obtain a silence-clipped performance audio. Normalize the audio of the silence-clipped performance audio to obtain a performance audio with a standard volume. Segment the notes in the performance audio with a standard volume to obtain the starting points of the performance notes. Convert the starting points of the performance notes into a spectrogram to obtain a performance spectrogram. Extract the pitch of the performance audio with a standard volume in a time series to obtain a performance pitch time series. Combine the performance spectrogram and the performance pitch time series to obtain the performance audio features.

10. A method and apparatus for generating a performance quality assessment driven by AI, characterized in that, Steps for implementing the performance quality evaluation generation method driven by any one of claims 1 to 9 include: An audio acquisition module for providing a repertoire for user selection and acquiring performance audio based on the user-selected repertoire. A feature acquisition module for extracting features from the performance audio to obtain performance audio features. A result acquisition module for evaluating the quality of the performance audio features to obtain an audio quality evaluation result. A visual feedback module for performing visual feedback according to the audio quality evaluation result.

Citation Information

Patent Citations

  • Scoring system for piano grading test

    CN107146497A

  • Intelligent scoring method for musical instrument

    CN111477249A

  • Piano playing learning method, electronic equipment and computer readable storage medium

    CN112598961A

  • Middle and primary school playing scoring method and system based on deep learning

    CN117746901A

  • System and method for music practice, and recording medium for recording program for realizing the same method

    JP1998187021A