Teaching method for characteristics of liuqin opera singing based on sound processing

By constructing a feature extraction network, adopting a differentiated framing strategy and multi-scale temporal convolution, and combining it with a consistency constraint loss function, the problem of insufficient accuracy in teaching Liuqin Opera singing was solved, achieving high-precision singing feature extraction and quantitative evaluation, and improving the scientificity and objectivity of teaching.

CN120822024BActive Publication Date: 2025-11-21SHANDONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511323642.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-11-21
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently capture the unique rhythmic system and vocal techniques of Liuqin Opera singing. Traditional acoustic feature separation methods are ineffective and lack quantitative evaluation. Furthermore, deep learning models struggle to incorporate domain knowledge, resulting in insufficient teaching accuracy.

Method used

A feature extraction network is constructed, employing a differentiated frame segmentation strategy and Fourier transform, combined with multi-scale temporal convolution and cross-channel attention mechanism, and a consistency constraint loss function is designed to achieve accurate extraction and quantitative evaluation of vocal features.

Benefits of technology

It achieves high-precision separation and quantitative evaluation of Liuqin Opera singing style, provides scientific teaching standards, and improves the objectivity and accuracy of singing quality assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822024B_ABST
    Figure CN120822024B_ABST
Patent Text Reader

Abstract

The application discloses a teaching method for the characteristics of the singing tune of the Liuqin opera based on sound processing, and comprises the following steps: constructing a singing tune audio library, and marking the time axis characteristics by experts to form a first positive sample; carrying out human voice separation, band filtering and differential frame processing based on rhythm types on the sample to generate a time-spectrum graph as a third positive sample; extracting a singing tune characteristic sequence through a characteristic extraction network; finally, using the trained network to carry out multi-dimensional evaluation on the singing tune of students, calculating quantitative indexes such as overall similarity, rhythm deviation and time sequence deviation, and providing objective singing feedback and individualized correction guidance for students. The application realizes the change of the teaching of the Liuqin opera from traditional experience transmission to modern quantitative analysis, and significantly improves the teaching efficiency and precision.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of digital signal processing and opera teaching, and particularly relates to a Liuqin opera singing tune feature teaching method based on sound processing. BACKGROUND

[0002] As an important branch of Chinese traditional opera, Liuqin opera combines unique local characteristics and vocal skills in its singing tune art. Its teaching and inheritance mainly rely on the "oral and mental instruction" mode. This traditional method highly depends on the personal experience of teachers, and has strong subjectivity and lacks quantitative standards. In particular, for the "board eye" rhythm system and "glissando" and "vibrato" complex skills, scholars have difficulty obtaining objective and accurate feedback.

[0003] With the development of technology, existing research attempts to apply audio analysis technology to opera teaching, but mostly uses general acoustic features (such as MFCC, fundamental frequency, etc.), which cannot effectively capture the unique singing tune types and rhythm systems of Liuqin opera. At the same time, traditional methods have poor separation effect when processing mixed audio of opera vocals and accompaniment, and lack the ability to model fine-grained features of singing tunes, which cannot meet the accuracy requirements of professional teaching.

[0004] In addition, opera audio data labeling costs are high, and there are complex music theory constraint relationships between features, making it difficult for traditional deep learning models to incorporate domain knowledge, resulting in prediction results that do not conform to artistic laws. Therefore, there is an urgent need for a specialized teaching method that can achieve high-precision voice separation, adapt to changes in opera rhythm, incorporate domain knowledge, and provide quantitative evaluation, to break through the traditional teaching bottleneck and realize the scientific inheritance of Liuqin opera art. SUMMARY

[0005] The purpose of the present application is to provide a Liuqin opera singing tune feature teaching method based on sound processing. To this end, the technical solution adopted by the present application is as follows:

[0006] A Liuqin opera singing tune feature teaching method based on sound processing, comprising the following steps:

[0007] S1: Construct a singing tune audio library; according to a preset labeling rule, label the feature category of each time point of the singing tune audio, obtain a labeling matrix, and generate a first positive sample;

[0008] S2: Perform voice separation on the first positive sample to obtain a second positive sample; then, according to the rhythm type, implement a differential framing strategy and Fourier transform on the second positive sample to obtain a third positive sample;

[0009] S3: According to the feature category, classify and extract the singing tune features in the third positive sample through a pre-constructed feature extraction network;

[0010] S4, collect student aria audio, select a contrast sample from the aria audio library; based on the trained feature extraction network, respectively construct the to-be-evaluated aria feature sequence of the student aria audio and the standard aria feature sequence of the contrast sample, and then perform student singing quality evaluation based on the standard aria feature sequence and the to-be-evaluated aria feature sequence.

[0011] Further, the structure of the feature extraction network is as follows:

[0012] The feature extraction network takes a shared encoder as the front end, followed by multiple parallel task decoders; the shared encoder is composed of a two-dimensional convolution layer and four stacked dilated temporal convolution modules; each dilated temporal convolution module is composed of a dilated one-dimensional convolution layer, a gated activation unit and a residual connection;

[0013] A task decoder is set up for each feature category c; each task decoder is composed of a one-dimensional convolution layer and a cross-channel attention mechanism;

[0014] The input of the shared encoder is the third positive sample, and the output is a shared feature sequence; the input of each task decoder is the shared feature sequence, and the output is a task feature vector of the corresponding feature category;

[0015] The task feature vector is input into a linear projection layer and a sigmoid activation function to generate the existence probability of the feature category c at each time point:

[0016]

[0017] wherein, represents the existence probability of the feature category c at time point t; T is the number of time steps.

[0018] Further, the total loss function of the feature extraction network is designed as follows:

[0019] The first part is the main loss function, which adopts a weighted binary cross-entropy loss; each feature category c is assigned a weight , The value of is inversely proportional to the frequency of the appearance of the feature category c in the third positive sample; the main loss function The formula is as follows:

[0020]

[0021] wherein, C is the total number of feature categories, is the true label of the feature category c at time point t;

[0022] The second part is the consistency constraint loss function; a prior rule matrix is defined , representing the feature class relationship with the feature class consistency constraint loss function The formula is as follows:

[0023]

[0024] The third part is a smoothing loss function to measure the continuity of the aria features in time, and the formula is as follows:

[0025]

[0026] The total loss function is:

[0027]

[0028] wherein, , and are hyperparameters for balancing the importance of each loss.

[0029] Further, after the feature extraction network is trained, the third positive sample is input into the feature extraction network to obtain the existence probability of each feature class at each time point;

[0030] By setting a specific dynamic threshold for each feature class, the existence probability is converted into a binary aria feature sequence, and the aria feature sequence is represented as:

[0031]

[0032] represents the aria feature corresponding to the feature class c at time point t;

[0033] represents the aria feature corresponding to the feature class c at time point t;

[0034] The label matrix is converted into a binary matrix, and the dimension of the binary matrix is exactly the same as that of the aria feature sequence;

[0035] The Euclidean distance between the aria feature sequence corresponding to the third positive sample and the binary matrix corresponding to the third positive sample is calculated, and when the Euclidean distance exceeds the set threshold, the model of the feature extraction network and the hyperparameters are adjusted for retraining.

[0036] Further, the student's student aria audio and the aria audio library use the same acquisition method, and the student aria audio is used as a first negative sample; the first negative sample is converted into a second negative sample by using the same steps as in step S2; the second negative sample is converted into a third negative sample; inputting the third negative sample into the trained feature extraction network to generate a to-be-evaluated aria feature sequence;

[0037] selecting the best version corresponding to the student aria audio from the aria audio library as a comparison sample, and dividing the comparison sample into a third positive sample by step S2, and inputting the third positive sample into the trained feature extraction network to generate a standard aria feature sequence.

[0038] Further, the process of evaluating the quality of student singing is as follows:

[0039] calculating the local distance between the first frame of the to-be-evaluated aria feature sequence and the first frame of the standard aria feature sequence ;

[0040] constructing a cumulative cost matrix ; representing the shortest distance from the starting point to the target point ; the construction process of the cumulative cost matrix is as follows:

[0041] firstly initializing ; then calculate the cumulative distance using the following recursive relationship:

[0042]

[0043] wherein, min represents the minimum value function;

[0044] In the process of constructing the cumulative cost matrix, the shortest distance from the starting point to the end point and the optimal path corresponding to the shortest distance are obtained; the optimal path is represented as:

[0045]

[0046] wherein, k represents the k path nodes, ;

[0047] calculate the overall similarity score , the formula is as follows:

[0048]

[0049] wherein, max represents the maximum value function; represents the theoretical maximum local distance, which is deduced from a local distance calculation formula; K is the length of the optimal path;

[0050] Calculate the rhythm deviation degree , the formula is as follows:

[0051]

[0052] Calculate the timing deviation degree , the timing deviation degree is the average offset between the optimal path and the diagonal line of the cumulative cost matrix, and the formula is as follows:

[0053]

[0054] Guide the student's Liuqin opera singing through the overall similarity score, the rhythm deviation degree and the timing deviation degree.

[0055] Further, the formula of the local distance is as follows:

[0056]

[0057] , wherein, is the weight of the feature category c, which is determined by training the main loss function in the feature extraction network, and the sum of the weights of different feature categories is 1; is an indicator function, when , the indicator function is 1, otherwise it is 0; represents that the tth time point of the singing feature sequence to be evaluated is a feature category c; represents that the tth time point of the standard singing feature sequence is a feature category c.

[0058] Compared with the prior art, the advantages of the present application are:

[0059] The present application first constructs a singing feature quantitative evaluation system based on a multi-dimensional dynamic time warping algorithm. By accurately comparing the student's singing with the standard template at the singing feature sequence level, multiple objective scores are generated, changing the evaluation method of the traditional oral teaching mode which relies on the subjective feeling of the teacher, and providing a scientific measurement standard for teaching.

[0060] The feature extraction network designed in the present application is not a general model, but deeply integrates the domain knowledge of Liuqin opera. By introducing a consistency constraint loss function, it ensures that the model prediction result conforms to basic music theory such as plate eye mutual exclusion; the differential framing strategy and multi-scale timing convolution adopted can adaptively capture features of different time scales from the instantaneous rhythm of flash plate to the long and sweet taste of the singing, and the extraction accuracy is much higher than that of traditional acoustic feature methods. BRIEF DESCRIPTION OF DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description only constitute some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0062] Figure 1 The flow chart of step S2 of the present application;

[0063] Figure 2 The flow chart of step S3 of the present application;

[0064] Figure 3 The flow chart of step S4 of the present application. DETAILED DESCRIPTION

[0065] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0066] The present embodiment is a sound processing-based teaching method for the singing aria features of the Liuqin Opera. Taking the classic singing sections of the Liuqin Opera as an example, the method is applied to the singing aria teaching scene in professional colleges and universities, and includes the following steps:

[0067] S1: Construct a singing aria audio library, and perform feature labeling on the singing aria audio library according to the time axis to generate a first positive sample;

[0068] The collection range of the singing aria audio library covers the two major schools of the North and the South, and contains classic plays such as “Fighting the Wild Ship” and “Drinking Mianye”. Each school has at least 50 complete singing sections, each with a duration of ≥3 minutes, and contains 12 typical singing aria types such as baby tune and crying aria. During the collection process, a SHURE SM7B condenser microphone is used to collect standard singing arias and singing audio of Liuqin Opera inheritors at a sampling rate of 48kHz and a bit depth of 24bit. The microphone is 0.5m away from the sound source, and the gain is set to 40dB.

[0069] The core rhythm features of the Liuqin Opera are the plate eye system, such as one plate one eye and one plate three eyes. Therefore, after obtaining the singing aria audio library, preset plate eye labeling rules are defined, and the start and end time labeling rules of singing aria skills including glissando and vibrato are defined.

[0070] The feature annotation matrix is obtained by annotating the singing voice audio library according to the annotation rules; the feature annotation process is to annotate a plurality of feature categories for each time point of the audio; the feature categories include but are not limited to {“board”, “eye”, “up slide”, “down slide”, “weak tremor”, “medium tremor”, “strong tremor”}.

[0071] In this embodiment, a labeling group composed of two first-class actors and one acoustic engineer performs triple annotation on all singing voice audios, and the consistency is verified by Cohen's Kappa coefficient, and the Cohen's Kappa coefficient is greater than 0.85.

[0072] The first positive sample is obtained after the feature annotation.

[0073] S2: performing a human voice separation operation on the first positive sample to obtain a second positive sample; then, a differential framing strategy and Fourier transform are performed on the second positive sample according to the rhythm type to obtain a third positive sample; referring to Figure 1 ;

[0074] In this embodiment, the Spleeter deep learning model (2-stem separation mode) is used to separate the human voice and the accompaniment. The deep learning model decomposes the input positive sample into human voice and accompaniment two independent tracks through a pre-trained convolutional neural network. For the high-frequency noise of the banhu and the lute in the lute opera accompaniment, the deep learning model inputs the WAV format audio with a sampling rate of 44.1Hz and a depth of 16bit, and outputs the separated double-channel file; the double-channel file includes a human voice track and an accompaniment track.

[0075] A Butterworth band-pass filter with a center frequency of 3kHz and a bandwidth of 1kHz is used to purify the human voice track twice after separation, to suppress the high-frequency noise of the instruments that are not completely separated, to obtain a second positive sample.

[0076] After filtering, the filtering effect is verified by the signal-to-noise ratio, which requires the signal-to-noise ratio of the human voice track to be greater than 24dB, to ensure the integrity of the specific human voice features such as “bright sound” and “brain sound” in subsequent feature extraction.

[0077] Two differential framing strategies are used for the second positive sample; the differential framing strategies include a fast section framing strategy and a slow section framing strategy.

[0078] The automatic switching of the framing strategy is realized by real-time detection of the short-time energy and zero-crossing rate of the audio, to ensure that the framing strategy is automatically matched with the singing rhythm.

[0079] In the embodiment, when the short-time energy of 3 continuous frames is greater than -20 dB and the zero-crossing rate is greater than 100 times per second, it indicates that the energy is high and the frequency changes fast, and then the fast tempo section is determined; when the short-time energy of 5 continuous frames is less than -30 dB and the zero-crossing rate is less than 50 times per second, it indicates that the energy is low and the frequency changes slowly, and then the slow tempo section is determined.

[0080] Both of the two differentiated framing strategies adopt a Hanning window function for smoothing processing to reduce inter-frame interference. The formula of the Hanning window function is as follows:

[0081]

[0082] wherein N represents the frame length; ;

[0083] In the embodiment, the framing strategy of the fast tempo section is as follows: the frame length is 256 points, corresponding to a time of 10.7 ms (based on a sampling rate of 44.1 Hz); the frame shift is 64 points, corresponding to a time of 2.7 ms (based on a sampling rate of 44.1 Hz); and the frame overlap rate is 75%. Fourier transform is performed on the second positive sample after framing, with the Fourier transform point number being 512, to generate a time-frequency spectrum graph as a third positive sample. The energy change trajectory of fast rhythm such as “flash plate” and “pile plate” is captured.

[0084] In the embodiment, the framing strategy of the slow tempo section is as follows: the frame length is 1024 points, corresponding to a time of 42.7 ms (based on a sampling rate of 44.1 Hz); the frame shift is 128 points, corresponding to a time of 5.3 ms (based on a sampling rate of 44.1 Hz); and the frame overlap rate is 87.5%. Fourier transform is performed on the second positive sample after framing, with the Fourier transform point number being 2048, to generate a time-frequency spectrum graph as a third positive sample. The fundamental frequency change of “glissando” in the slow tempo section can be captured.

[0085] S3: according to the feature category, a pre-constructed feature extraction network is used to classify and extract the singing section features in the third positive sample.

[0086] The feature extraction network has a shared encoder as a front end, followed by multiple parallel task decoders. Referring to Figure 2 ;

[0087] The shared encoder takes the third positive sample as input. First, a two-dimensional convolution layer is used for preliminary down-sampling and feature abstraction; the convolution kernel of the two-dimensional convolution layer is , and the step size is Subsequently, four stacked dilated temporal convolutional modules are connected. Each dilated temporal convolutional module consists of a dilated one-dimensional convolutional layer, a gated activation unit, and residual connections. The convolutional kernel of the dilated one-dimensional convolutional layer is 3, and the dilation factor increases exponentially layer by layer (e.g., 1, 2, 4, 8) to exponentially expand the receptive field and capture multi-scale temporal information from short-term audio time (e.g., "board") to long-term context-dependent information (e.g., "long melody"). The gated activation unit uses a GLU (Gated Linear Unit) instead of the traditional ReLU to dynamically modulate the information flow and enhance the feature extraction network's ability to select important features. The residual connections are used to alleviate the gradient vanishing problem of the feature extraction network. The output of the shared encoder is a shared feature sequence. , where T is the number of time steps.

[0088] For each feature category A task decoder is established. Each task decoder is essentially a one-dimensional convolutional layer followed by a cross-channel attention mechanism. The cross-channel attention mechanism automatically learns and emphasizes the channel dimension information most relevant to the corresponding feature category of the decoder among the shared features. The calculation formula is as follows:

[0089]

[0090]

[0091]

[0092] in, and Indicate feature category Learnable parameters Indicate feature category Attention weights; Indicate feature category Projection sharing features; Indicate feature category Task characteristics; for function, It is the hyperbolic tangent function.

[0093] Finally, the task features are input into a linear projection layer and a sigmoid activation function to generate the existence probability of the feature category c at each time point. , This represents the probability of feature category c existing at time point t.

[0094] The feature extraction network learns the extraction task for all feature categories simultaneously. The total loss function consists of three parts; the considerations and design of each part are as follows:

[0095] The first part is the main loss function; a weighted binary cross-entropy loss is adopted. Since the number of samples in the third positive sample is extremely unbalanced (for example, the number of "with board" and "without board"), a weight is assigned to each feature category c ; The value of the weight is inversely proportional to the frequency of the occurrence of the feature category c in the third positive sample; the main loss function is as follows:

[0096]

[0097] where C is the total number of feature categories, is the true label of the feature category c at time point t.

[0098] The second part is the consistency constraint loss function, which introduces constraints according to the opera knowledge of Liuqin opera; by defining a priori rule matrix , represents the relationship between the feature category and the feature category ; for example, -1 represents negative correlation, +1 represents positive correlation, and 0 represents no constraint. In actual Liuqin opera clips, "board" and "eye" are usually mutually exclusive, and "strong tremor" is unlikely to appear near the starting point of "glissando" skill. The consistency constraint loss function is as follows:

[0099]

[0100] The third part is the smoothing loss function, which considers that the singing feature has continuity in time, for example, "tremolo" usually lasts for a period of time, so a first-order difference penalty is used to punish the sharp fluctuation of the existence probability, so that the output existence probability sequence is smoother. The smoothing loss function is as follows:

[0101]

[0102] The total loss function is:

[0103]

[0104] where , and are hyperparameters for balancing the importance of each loss. , and The values of the above parameters are determined according to the actual needs of the feature extraction network, and can be generally set to 0.6, 0.2, and 0.2.

[0105] After the feature extraction network is trained, the third positive sample input value is input into the feature extraction network to obtain the existence probability of each feature category at each time point. By setting a specific judgment threshold for each feature category, the probability value is converted into a binary aria feature sequence; for example, the judgment threshold of the "board" feature is set to 0.9.

[0106] The aria feature sequence is represented as:

[0107]

[0108] represents that at time point t, there is an aria feature corresponding to feature category c;

[0109] represents that at time point t, there is no aria feature corresponding to feature category c.

[0110] The labeling matrix is converted into a binary matrix, and the dimensions of the binary matrix are completely consistent with the aria feature sequence. In this embodiment, when feature category c is labeled at time point t, the corresponding element of the binary matrix is set to 1, otherwise it is set to 0;

[0111] The Euclidean distance between the aria feature sequence corresponding to the third positive sample and the binary matrix corresponding to the third positive sample is calculated. When the Euclidean distance exceeds the set threshold, the model and hyperparameters of the feature extraction network need to be adjusted and retrained.

[0112] In this embodiment, when the Euclidean distance exceeds 0.05, the training is restarted.

[0113] S4, collect student aria audio, and use the trained feature extraction network to evaluate and correct the student aria audio. Refer to Figure 3 ;

[0114] In this embodiment, the student aria audio and the aria audio library use the same collection method, and the student aria audio is used as the first negative sample; the same steps as S2 are used to convert the first negative sample into a second negative sample, and the differential framing strategy is used to convert the second negative sample into third negative samples; and the third negative samples are input into the trained feature extraction network to generate aria feature sequences to be evaluated.

[0115] The best version corresponding to the student aria audio is selected from the aria audio library as a comparison sample, and the comparison sample is framed into third positive samples by step S2, and the third positive samples are input into the trained feature extraction network to generate A standard vocal technique feature sequence.

[0116] Calculate the first of the vocal feature sequences to be evaluated The first frame and the first of the standard vocal feature sequence Local distance per frame The local distance reflects the degree of difference in all vocal features between the two time points.

[0117] In this embodiment, the local distance is calculated using the weighted Hamming distance, as shown in the following formula:

[0118]

[0119] in, The weights of feature category c are determined through training the main loss function in the feature extraction network, and the sum of the weights of different feature categories is 1. Let c be the indicator function. When the t-th time point of the vocal feature sequence to be evaluated is feature category c, it is represented as: When the t-th time point of the standard singing style feature sequence is feature category c, it is represented as: ;when The indicator function is 1 if it is true, and 0 otherwise.

[0120] Constructing the cumulative cost matrix , , ; Indicates starting from the point To the target point The shortest distance; the construction process of the cumulative cost matrix is ​​as follows:

[0121] First initialize , The first element of the vocal feature sequence to be evaluated is... The local distance between each frame and the first frame of the standard vocal feature sequence;

[0122] Then, the cumulative distance is calculated using the following recursive relation:

[0123]

[0124] Where min represents the minimum value function.

[0125] During the construction of the cumulative cost matrix, the starting point can be obtained. To the finish line The shortest distance and the optimal path corresponding to the shortest distance.

[0126] Optimal path Represented as:

[0127]

[0128] wherein, represents k path nodes, , ;

[0129] Let the length of the optimal path be K;

[0130] Calculate the overall similarity score , the formula is as follows:

[0131]

[0132] wherein, max represents the maximum value function; represents the theoretical maximum local distance, which is inferred to be 1 from the local distance calculation formula.

[0133] Calculate the rhythm deviation degree , the formula is as follows:

[0134]

[0135] The timing deviation degree can also be calculated from the optimal path. Calculate the average offset between the optimal path and the diagonal line of complete alignment, which is used to judge the overall drag or grab. The calculation formula is as follows:

[0136]

[0137] Guide students to play the Liuqin opera by using the overall similarity score, rhythm deviation degree and timing deviation degree.

[0138] The above formulas are all dimensionless forms, only numerical values are used for calculation. These formulas are based on a large amount of data and obtained through software simulation, aiming to be as close to the actual situation as possible. The preset parameters in the formula can be adjusted by those skilled in the art according to specific needs.

[0139] In the description of the present specification, the description of the terms "one embodiment", "example", "specific example" and the like means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0140] The preferred embodiments of the application disclosed above are only to facilitate the elucidation of the application. The preferred embodiments do not describe all the details of the application and limit the application to the specific embodiments. Obviously, many modifications and variations can be made in light of the teachings above. The description is chosen and described in order to provide the best illustration of the application and its practical application to those skilled in the art and to enable those skilled in the art to best utilize the application. The application is limited only by the claims and their full scope and equivalents.

Claims

1. A teaching method for the vocal characteristics of Liuqin Opera based on sound processing, characterized in that, Includes the following steps: S1: Construct a vocal music audio library; according to preset annotation rules, label each time point of the vocal music audio with preset feature categories to obtain an annotation matrix, and generate the first positive sample; S2: Perform voice separation on the first positive sample to obtain the second positive sample; then apply a differentiated framing strategy and Fourier transform to the second positive sample according to the rhythm type to obtain the third positive sample; S3: Based on the aforementioned feature categories, extract the vocal features from the third positive sample using a pre-constructed feature extraction network; S4. Collect student singing audio and select comparison samples from the singing audio library; construct the singing feature sequence to be evaluated and the standard singing feature sequence of the comparison sample based on the trained feature extraction network; and then evaluate the student's singing quality based on the standard singing feature sequence and the singing feature sequence to be evaluated. The structure of the feature extraction network is as follows: The feature extraction network uses a shared encoder as the front end, followed by multiple parallel task decoders; the shared encoder consists of a two-dimensional convolutional layer and four stacked dilated temporal convolutional modules; each dilated temporal convolutional module consists of a dilated one-dimensional convolutional layer, a gated activation unit, and residual connections; A task decoder is set up for each feature category c; each task decoder consists of a one-dimensional convolutional layer and a cross-channel attention mechanism. The input of the shared encoder is the third positive sample, and the output is a shared feature sequence; the input of each task decoder is the shared feature sequence, and the output is a task feature vector of the corresponding feature category. The task feature vector is input into a linear projection layer and a sigmoid activation function to generate the existence probability of the feature class c at each time point: in, This represents the probability of feature category c existing at time point t; T is the number of time steps.

2. The method according to claim 1, characterized in that, The total loss function of the feature extraction network is designed as follows: The first part is the main loss function, which uses a weighted binary cross-entropy loss; weights are assigned to each feature category c. , The value is inversely proportional to the frequency of feature category c in the third positive sample; the main loss function The formula is as follows: Where C is the total number of feature categories. The true label of feature category c at time point t; The second part is the consistency constraint loss function; defining the prior rule matrix. , Indicate feature category With feature category Relationship; Consistency constraint loss function The formula is as follows: The third part is the smoothing loss function, which measures the continuity of vocal characteristics over time. The formula is as follows: The total loss function is: in, , and This is a hyperparameter used to balance the importance of various losses.

3. The method according to claim 2, characterized in that, After the feature extraction network is trained, the third positive sample is input into the feature extraction network to obtain the existence probability of each feature category at each time point; By setting a specific dynamic threshold for each feature category, the existence probability is transformed into a binary sequence of vocal feature characteristics, which is represented as: This indicates that at time point t, there exists a singing style feature corresponding to feature category c; This indicates that at time point t, there is no singing style feature corresponding to feature category c; The annotation matrix is ​​converted into a binary matrix, and the dimension of the binary matrix is ​​exactly the same as that of the singing feature sequence; Calculate the Euclidean distance between the vocal feature sequence corresponding to the third positive sample and the binary matrix corresponding to the third positive sample. If the Euclidean distance exceeds the set threshold, adjust the model and hyperparameters of the feature extraction network and retrain.

4. The method according to claim 3, characterized in that, The student's singing audio and the singing audio library are acquired using the same method, with the student's singing audio serving as the first negative sample. The first negative sample is then converted into a second negative sample using the same steps as in step S2. Finally, the differentiated framing strategy is used to frame and convert the second negative sample into... A third negative sample is generated; the third negative sample is input into the trained feature extraction network to generate... One vocal feature sequence to be evaluated; The best version of the student's singing voice is selected from the singing voice audio library as a comparison sample, and the comparison sample is framed in step S2. A third positive sample, and the... A third positive sample is input into the trained feature extraction network to generate... A standard vocal technique feature sequence.

5. The method according to claim 4, characterized in that, The process for evaluating student singing quality is as follows: Calculate the first of the vocal feature sequences to be evaluated The first frame and the first of the standard vocal feature sequence Local distance per frame ; Constructing the cumulative cost matrix ; Indicates starting from the point To the target point The shortest distance; the process of constructing the cumulative cost matrix is ​​as follows: First initialize Then, the cumulative distance is calculated using the following recursive relation: in, Describes the minimum value function; In the process of constructing the cumulative cost matrix, the starting point is obtained. To the finish line The shortest distance and the optimal path corresponding to the shortest distance; the optimal path Represented as: in, Represents k path nodes, ; Calculate the overall similarity score The formula is as follows: in, Represents the maximum value function; This represents the theoretical maximum local distance, which is inferred from the local distance calculation formula; K The length of the optimal path; Calculate rhythm deviation The formula is as follows: Calculate the time series deviation The temporal deviation is the average offset between the optimal path and the diagonal of the cumulative cost matrix, as shown in the following formula: The overall similarity score, rhythm deviation, and temporal deviation are used to guide students in singing Liuqin Opera.

6. The method according to claim 5, characterized in that, The formula for the local distance is as follows: in, The weights of feature category c are determined through training the main loss function in the feature extraction network, and the sum of the weights of different feature categories is 1. For indicator functions, when The indicator function is 1 if it is true, and 0 otherwise. The t-th time point of the vocal feature sequence to be evaluated is the feature category c; Let the t-th time point of the standard singing style feature sequence be the feature category c.

Citation Information

Patent Citations

  • Multi-scale multi-view-based Chinese opera singing style identification method

    CN111402919A

  • A singing segment multi-dimensional evaluation method and terminal based on deep learning

    CN119763611A