A method for visual analysis of sound in vocal music teaching
By using multi-directional microphone arrays and neck and throat muscle fiber sensors in vocal teaching, combining machine learning and time series analysis to generate intuitive visual results, the problem of neglecting physiological factors and difficult to intuitively understand the results of analysis in the prior art is solved, and in-depth evaluation of vocal level and the provision of personalized training plans are achieved.
Patent Information
- Application Number
- CN202510269415.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing vocal teaching technology mainly focuses on sound signal analysis, ignores physiological factors closely related to sound production, resulting in insufficient in-depth understanding and comprehensive evaluation of students' singing skills, and the analysis results are difficult to understand intuitively, and lack real-time feedback and personalized guidance.
A multi-directional microphone array and neck and throat muscle fiber sensor are used to collect sound signals and physiological data. Through machine learning algorithms and time series analysis, combined with multi-dimensional correlation analysis, intuitive visual results are generated, including real-time data dashboards and historical data and predicted trend comparison charts.
It realizes accurate collection and in-depth analysis of sound signals and physiological data, can accurately evaluate vocal level and existing problems, provide personalized training plans, improve vocal skills and performance levels, and facilitate the application of teaching and training practice through intuitive visualization results.
Smart Images

Figure CN119785744B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of music education. More specifically, the present invention relates to a method for visual analysis of sound in vocal music teaching. Background Art
[0002] As an important branch in the field of music education, vocal music teaching aims to improve students' sound control ability, singing skills, and musical expressiveness through systematic training and guidance. With the continuous development of music education and the continuous progress of technology, the field of vocal music teaching is also constantly exploring and applying new technical means to improve teaching effects and meet the personalized needs of students.
[0003] In current vocal music teaching, the existing technical means mainly focus on the analysis and processing of sound signals. Specifically, these technologies usually collect students' singing voices through professional audio recording devices, and then use audio analysis software to perform detailed analysis and processing on the sound signals. This process includes multiple links such as filtering, noise reduction, spectrum analysis, and pitch detection of sound signals, aiming to extract key parameters related to singing performance, such as volume, pitch, timbre, and resonance effect. Based on these analysis results, teachers can provide targeted guidance and improvement on students' singing skills.
[0004] Although the existing vocal music teaching technologies have improved the objectivity and accuracy of teaching to a certain extent, there are still some obvious deficiencies. First of all, these technologies mainly focus on the analysis of sound signals and ignore the physiological factors closely related to sound production, such as the movement states of neck and laryngeal muscles, which limits the in-depth understanding and comprehensive evaluation of students' singing skills. Secondly, the analysis results of the existing technologies are usually presented in the form of numbers and charts, which is difficult for students without a music professional background to understand and is not conducive to the improvement of teaching effects. In addition, the existing technologies also lack real-time feedback and personalized guidance during students' singing processes and are difficult to meet the diverse learning needs of students. Therefore, the field of vocal music teaching still needs to explore more comprehensive, intuitive, and easy-to-understand teaching methods and technical means. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for visual analysis of sound in vocal music teaching, through the following solutions to solve the problems raised in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solution: A method for visual analysis of sound in vocal music teaching, including:
[0007] Step 1: Sensor setting: including the setting of a microphone array and neck and laryngeal muscle fiber sensors for sound collection and physiological data collection;
[0008] Step 2: Data collection: Used to collect voice data and physiological data;
[0009] Step 3: Voice data analysis: Used to analyze the voice data collected in Step 2 to obtain dynamic range and timbre change data, resonance and vocal efficiency data, and technique and adaptability data;
[0010] Step 4: Multidimensional correlation analysis: By combining machine learning algorithms and time series analysis, comprehensively analyze the correlation between voice data and neck and laryngeal muscle movement data;
[0011] Step 5: Visualization: Includes a real-time data dashboard and a comparison chart of historical data and predicted trends, used to convert data into an intuitive form for display.
[0012] Preferably, the voice collection uses a multi-directional microphone array, which is arranged around the singer on a semi-circular track, about 0.8 - 1.2 meters away from the singer, and is equipped with a professional audio interface to convert the voice signal into a digital signal and transmit it to the computer.
[0013] Preferably, the physiological data is collected by winding fiber optic sensors around the muscle groups in the neck and larynx, used to sense muscle deformation and tension changes.
[0014] Preferably, the fiber optic sensor is a special fiber optic sensor made based on the fiber Bragg grating principle. The core component is a section of specially treated optical fiber, in which refractive index modulation regions are periodically distributed in the core to form a Bragg grating. The outer layer of the fiber optic sensor uses a biocompatible material that fits the human skin to ensure that it will not cause irritation or discomfort to the singer's skin during long-term wearing, and the internal optical fiber is encapsulated.
[0015] Preferably, the voice data is continuously collected by the microphone array at a sampling rate of 48 kHz and a quantization precision of 24 bits when the singer is practicing vocal music or singing different song segments. During the collection process, segment markers are made according to different stages of singing for subsequent targeted analysis.
[0016] Preferably, the physiological data records the muscle deformation and tension data through the neck and laryngeal muscle fiber optic sensors at a sampling rate of 1000 Hz.
[0017] Preferably, the dynamic range and timbre change data include the volume attenuation ratio from the lowest note to the highest note, the timbre brightness index, the timbre consistency during the range span, and the vibrato rate and amplitude ratio. The resonance and vocal efficiency data include the head cavity resonance ratio, the vocal cord closure degree during vocalization, the breath flow rate and pitch stability coefficient, and the resonance cavity conversion flexibility. The technique and adaptability data include the grace note accuracy, the rhythm flexibility index, the voice intensity control precision, and the voice fatigue index.
[0018] Preferably, the analysis method for the volume attenuation ratio from the lowest note to the highest note is specifically as follows: perform time-domain analysis on the sound signal, identify the start and end positions of the lowest note and the highest note, calculate the root mean square value of the volume between the two, and obtain the volume attenuation ratio through the formula where is the root mean square value of the volume of the lowest note, is the root mean square value of the volume of the highest note. The analysis method for the timbre brightness index is specifically as follows: use Fourier transform to convert the sound signal to the frequency domain, assign weights to different frequency bands according to the perception characteristics of the human ear for sounds of different frequencies, and calculate the timbre brightness index through the formula where is the weight of the k-th frequency band, is the energy proportion of this frequency band, m represents the number of divided frequency bands. The analysis method for the timbre consistency during the range span is specifically as follows: at the key transition points of the range span, extract the Mel-frequency cepstral coefficient feature vectors of the sound, calculate the cosine similarity between adjacent feature vectors, and obtain the timbre consistency index using the formula where and are respectively the i-th MFCC coefficients of two adjacent transition points, n is the dimension of the MFCC coefficients. The analysis method for the vibrato rate and amplitude ratio is specifically as follows: extract the instantaneous frequency of the sound signal through Hilbert transform, detect the vibrato part, calculate the number of frequency fluctuations per unit time to obtain the vibrato rate, and use the ratio of the maximum amplitude of the frequency fluctuation to the average frequency as the vibrato amplitude ratio, that is where where represents the vibrato rate, represents the number of frequency fluctuations, represents time, represents the amplitude ratio, represents the maximum frequency fluctuation amplitude, represents the average frequency.
[0019] Preferably, the analysis method for the head cavity resonance ratio is specifically as follows: based on the formant analysis of the sound signal, use linear predictive coding technology to determine the frequency and intensity of the formants, and obtain the head cavity resonance ratio by comparing the energy of the head cavity resonance formants with the total sound energy, using the formula Calculate the proportion of head cavity resonance, represent the energy of the resonance peaks related to head cavity resonance, represent the total energy of the sound signal. The analysis method of vocal cord closure degree during sound production is as follows: Analyze the harmonic structure of the sound signal, calculate the amplitude ratio and phase relationship between the harmonics and the fundamental wave, and estimate the vocal cord closure degree using an empirical formula , and represent the amplitude and phase of the j-th harmonic, and represent the amplitude and phase of the fundamental wave, and represent empirical coefficients, p represents the number of harmonics. The analysis method of the breath flow rate and pitch stability coefficient is as follows: Combine the breath flow rate data collected by the breath sensor and the pitch change of the sound signal, and calculate the ratio of the standard deviation of the pitch to the average breath flow rate. The formula is , represents the standard deviation of the pitch, represents the average gas flow rate. The analysis method of the flexibility of resonance cavity conversion is as follows: By monitoring the rapid change of the energy of the sound signal in different frequency bands and the migration of the resonance peaks, calculate the rate of change of the resonance peak frequency and intensity in adjacent time periods as an index of the flexibility of resonance cavity conversion. The formula , and are the resonance peak frequency and intensity at the t-th moment respectively, q represents the total number of time periods for analysis, represents the flexibility of resonance cavity conversion, represents the resonance peak frequency at the (t + 1)-th moment, represents the intensity at the (t + 1)-th moment.
[0020] Preferably, the analysis method of the accuracy of the grace notes is as follows: Compare the grace note part of the performance with the grace notes in the standard musical score, and calculate the errors in pitch, duration, and start time between the two. Use the formula to calculate the accuracy index, , and represent the pitch, duration, and start time of the standard grace notes respectively, , and represent the pitch, duration, and start time of the actually performed grace notes respectively, , and represent the maximum values of the pitch, duration, and start time of all grace note parameters respectively, , and respectively represent the minimum values of pitch, duration, and start time of all ornament parameters, s represents the number of ornaments, and the analysis method of the rhythm flexibility index is specifically as follows: Use the beat tracking algorithm to determine the actual rhythm of the singing, compare it with the standard rhythm, calculate the statistic of the rhythm deviation within a certain time window, and use the formula to calculate the rhythm flexibility index, represents the variance of the rhythm deviation. The analysis method of the voice intensity control accuracy is specifically as follows: Analyze the dynamic range of the voice signal, divide it into multiple levels, calculate the accuracy and fineness of the intensity conversion between different levels during actual singing, and use the formula to calculate, represents whether the u-th intensity conversion is accurate, 1 for accurate and 0 for inaccurate, v represents the total number of intensity conversions. The analysis method of the voice fatigue index is specifically as follows: Monitor the long-term change trend of the voice signal, including the drift of the fundamental frequency and the attenuation of harmonics. By establishing a voice fatigue model, use the formula to calculate the voice fatigue index, represents the change amount of the fundamental frequency, and respectively represent the change amounts of the amplitude and phase of the w-th harmonic, represents the amplitude of the w-th harmonic at the initial moment, , and are model coefficients, and x is the total number of harmonics.
[0021] Preferably, the machine learning algorithm sets the voice feature vector as , the neck and laryngeal muscle movement feature vector as , and the prediction function obtained through training as , represents the weight corresponding to the voice feature, represents the weight corresponding to the neck and laryngeal muscle movement feature, b is the bias term, n1 represents the number of elements in the voice feature vector, s i1 represents the i1-th feature in the voice feature vector, k1 represents the number of elements in the neck and laryngeal muscle movement feature vector, m j1 represents the j1-th feature in the neck and laryngeal muscle movement feature vector.
[0022] Preferably, the time series analysis establishes an ARIMA(p, d, q) model based on the voice fatigue index F t and any parameter of the neck and laryngeal muscle movement. Let F t after d times of differencing be , and establish the model , represents the autoregressive coefficient, represents the moving average, denotes the white noise term, p1 denotes the autoregressive order, i2 is used to iterate through each lag order from 1 to p1 in the autoregressive part, q1 denotes the moving average order, j2 is used to iterate through each lag order from 1 to q1 in the moving average part, and t1 denotes the current time.
[0023] Preferably, the real-time data dashboard visualizes the metrics output in steps 3 and 4. Each metric is represented by a circular dial, and the pointer points to the current numerical position in real time. According to the set threshold range on the dial, different colored regions are divided. Green represents the normal range, yellow represents approaching the warning range, and red represents exceeding the warning range.
[0024] Preferably, the historical data and predicted trend comparison chart shows the historical data of the singer and the predicted trend of the multi-dimensional correlation model for the future by drawing a comprehensive chart. The horizontal axis of the chart represents time, and the vertical axis represents the values of various metrics. Different colored lines are used to represent the change trends of the historical data and the future trends predicted by the model respectively. The blue line represents the actual data change of the sound resonance effect metric in the past month, and the red line represents the development trend of this metric in the next two weeks based on the model prediction. On the chart, annotations and notes are added to facilitate users to understand the data changes. For possible abnormal situations in the predicted trend, early warning annotations are made. When it is predicted that the voice fatigue index will exceed the warning threshold in the next week, a yellow triangle is marked on the chart, and a prompt is given "Please note that the voice fatigue may reach a dangerous level in the next week. Please adjust the training plan", and a detailed analysis is carried out in a graphical way. Specifically, through a tree diagram, the contribution degree of each factor in the voice characteristics and neck and laryngeal muscle movement parameters to the abnormal metric is shown, and different colors and line thicknesses are used to represent the influence degree. Clicking on each factor in the tree diagram can also display detailed historical data.
[0025] The technical effects and advantages of the present invention:
[0026] The present invention realizes the precise acquisition of sound signals and physiological data by setting up a multi-directional microphone array and fiber optic sensors for the neck and laryngeal muscles. The microphone array can capture the singer's voice omnidirectionally, while the fiber optic sensors can accurately sense the deformation and tension of the neck and laryngeal muscles, providing a high-quality data basis for subsequent analysis. In the data acquisition stage, the system continuously records sound and physiological data at a high sampling rate and marks them in segments according to different stages of singing. At the same time, metadata such as the song name and singer information are also recorded to facilitate the classification management and in-depth analysis of subsequent data. This step ensures the comprehensiveness and accuracy of the data. The analysis of sound data is the core link of this method. By performing multi-dimensional analysis on the collected sound data, including aspects such as dynamic range, timbre change, resonance effect, vocal efficiency, vocal techniques, and adaptability, the system can accurately evaluate the singer's vocal level and existing problems, providing a scientific basis for subsequent personalized teaching and training. By introducing machine learning algorithms and time series analysis, the correlation between sound characteristics and neck and laryngeal muscle movements is deeply explored. This can not only reveal the mystery of the vocalization mechanism but also predict the possible future change trends of sound indicators. Based on these analysis results, teachers can formulate personalized training plans for singers to improve their vocal techniques and performance levels. The analysis results are presented in an intuitive and easy-to-understand manner. Through a real-time data dashboard and a comparison chart of historical data and predicted trends, singers and teachers can clearly see the values and change trends of various indicators. At the same time, the system can also provide warnings and detailed analysis windows to help singers discover problems in a timely manner and adjust their training plans. This visual presentation method makes the analysis results more intuitive and easy to understand, facilitating their application in teaching and training practices. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic diagram of the overall structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] Refer to Figure 1 A method for sound visualization analysis in vocal music teaching shown, the specific steps include:
[0030] Step 1: Sensor setting: including the setting of a microphone array and fiber optic sensors for the neck and laryngeal muscles, for sound acquisition and physiological data acquisition.
[0031] The sound collection uses a multi-directional microphone array, which is arranged around the singer on a semi-circular track, about 0.8 - 1.2 meters away from the singer, and is equipped with a professional audio interface to convert the sound signal into a digital signal and transmit it to the computer.
[0032] The physiological data collection is achieved by winding fiber optic sensors around the muscle groups in the neck and larynx to sense muscle deformation and tension changes.
[0033] The fiber optic sensor is specifically a special fiber optic sensor made based on the fiber Bragg grating principle. The core component is a section of specially treated optical fiber, in which refractive index modulation regions are periodically distributed in the core to form a Bragg grating. The outer layer of the fiber optic sensor uses a biocompatible material that conforms to the human skin to ensure that it will not cause irritation or discomfort to the singer's skin during long-term wearing. The internal optical fiber is encapsulated, which can not only ensure high-sensitivity perception of small muscle deformations and tension changes, but also has good flexibility to adapt to the complex movements of the neck and larynx muscles.
[0034] Compared with traditional electromyography sensors, the fiber optic sensor has significant advantages. Since it is based on optical signal transmission, it is extremely resistant to electromagnetic interference and can work stably in a complex electronic device environment. Moreover, it can measure the dynamic changes of muscles with extremely high precision, and the deformation can be accurate to the micron level. For example, when the laryngeal muscles produce extremely subtle contractions or relaxations during vocalization, the period of the Bragg grating in the fiber optic sensor will change accordingly, resulting in a change in the wavelength of the reflected light. Through a high-precision wavelength demodulation device, this wavelength change can be accurately converted into muscle deformation and tension data, providing extremely reliable and accurate data support for subsequent analysis of muscle movements during vocalization.
[0035] Step 2: Data collection: used to collect sound data and physiological data.
[0036] The sound data is continuously collected by the microphone array at a sampling rate of 48 kHz and a quantization precision of 24 bits when the singer is practicing vocals or singing different song segments. During the collection process, segment markers are made according to different stages of the performance for subsequent targeted analysis.
[0037] The physiological data records the muscle deformation and tension data through the fiber optic sensors in the neck and larynx muscles at a sampling rate of 1000 Hz.
[0038] Step 2 also records metadata such as the name of the song being sung, singer information, and singing time, which facilitates the classification and management of the data.
[0039] Step 3: Voice Data Analysis: Used to analyze the voice data collected in Step 2 to obtain dynamic range and timbre change data, resonance and vocal efficiency data, as well as technique and adaptability data.
[0040] The dynamic range and timbre change data include the volume attenuation ratio from the lowest note to the highest note, the timbre brightness index, the timbre consistency during the range crossing, and the vibrato rate and amplitude ratio. The resonance and vocal efficiency data include the head cavity resonance ratio, the vocal cord closure degree during vocalization, the breath flow rate and pitch stability coefficient, and the resonance cavity conversion flexibility. The technique and adaptability data include the grace note accuracy, the rhythm flexibility index, the voice intensity control precision, and the voice fatigue index.
[0041] The specific analysis method for the volume attenuation ratio from the lowest note to the highest note is as follows: Perform time-domain analysis on the voice signal, identify the start and end positions of the lowest note and the highest note, calculate the root mean square value of the volume between the two, and obtain the volume attenuation ratio through the formula where is the root mean square value of the volume of the lowest note, is the root mean square value of the volume of the highest note. The specific analysis method for the timbre brightness index is as follows: Use Fourier transform to convert the voice signal to the frequency domain, assign weights to different frequency bands according to the perception characteristics of the human ear for sounds of different frequencies, and calculate the timbre brightness index through the formula where is the weight of the k-th frequency band, is the energy proportion of this frequency band, and m represents the number of divided frequency bands. The specific analysis method for the timbre consistency during the range crossing is as follows: At the key transition points of the range crossing, extract the Mel-frequency cepstral coefficient feature vectors of the voice, calculate the cosine similarity between adjacent feature vectors, and obtain the timbre consistency index using the formula where and are the i-th MFCC coefficients of two adjacent transition points respectively, and n is the dimension of the MFCC coefficients. The specific analysis method for the vibrato rate and amplitude ratio is as follows: Extract the instantaneous frequency of the voice signal through Hilbert transform, detect the vibrato part, calculate the number of frequency fluctuations per unit time to obtain the vibrato rate, and the ratio of the maximum amplitude of the frequency fluctuation to the average frequency as the vibrato amplitude ratio, that is , , represents the vibrato rate, represents the number of frequency fluctuations, represents the time, represents the amplitude ratio, represents the maximum frequency fluctuation amplitude, represents the average frequency.
[0042] The analysis method of the head cavity resonance ratio is specifically as follows: Based on the formant analysis of the sound signal, the linear prediction coding technology is used to determine the frequency and intensity of the formants. By comparing the energy of the head cavity resonance formants with the total sound energy, the formula is used to calculate the head cavity resonance ratio. represents the energy of the resonance formants related to the head cavity resonance. represents the total energy of the sound signal. The analysis method of the vocal cord closure degree during vocalization is specifically as follows: Analyze the harmonic structure of the sound signal. By calculating the amplitude ratio and phase relationship between the harmonics and the fundamental wave, the vocal cord closure degree is estimated using the empirical formula , and represent the amplitude and phase of the j-th harmonic. and represent the amplitude and phase of the fundamental wave. and represent the empirical coefficients, p represents the number of harmonics. The analysis method of the breath flow rate and pitch stability coefficient is specifically as follows: Combining the breath flow rate data collected by the breath sensor and the pitch change of the sound signal, calculate the ratio of the standard deviation of the pitch to the average breath flow rate. The formula is , represents the standard deviation of the pitch. represents the average gas flow rate. The analysis method of the resonance cavity conversion flexibility is specifically as follows: By monitoring the rapid change of the sound signal energy in different frequency bands and the migration of the formants, calculate the rate of change of the formant frequency and intensity in adjacent time periods as an index of the resonance cavity conversion flexibility. The formula , and are the formant frequency and intensity at the t-th moment respectively, q represents the total number of analysis time periods. represents the resonance cavity conversion flexibility. represents the formant frequency at the (t + 1)-th moment. represents the intensity at the (t + 1)-th moment.
[0043] The analysis method of the grace note accuracy is specifically as follows: Compare the grace note part sung with the grace notes in the standard music score. By calculating the errors in pitch, duration, and start time between the two, use the formula to calculate the accuracy index. , and represent the pitch, duration, and start time of the standard grace note respectively. , and represent the pitch, duration, and start time of the actually sung grace note respectively. , and respectively represent the maximum values of the pitch, duration, and start time of all ornament parameters, , and respectively represent the minimum values of the pitch, duration, and start time of all ornament parameters. s represents the number of ornaments. The specific analysis method of the rhythm flexibility index is as follows: Use the beat tracking algorithm to determine the actual rhythm of the performance, compare it with the standard rhythm, calculate the statistic of the rhythm deviation within a certain time window, and calculate the rhythm flexibility index through the formula Calculate the rhythm flexibility index. represents the variance of the rhythm deviation. The specific analysis method of the voice intensity control accuracy is as follows: Analyze the dynamic range of the voice signal, divide it into multiple levels, calculate the accuracy and fineness of the intensity conversion between different levels during the actual performance, and calculate through the formula Calculate. represents whether the u-th intensity conversion is accurate, 1 for accurate and 0 for inaccurate. v represents the total number of intensity conversions. The specific analysis method of the voice fatigue index is as follows: Monitor the long-term change trend of the voice signal, including the drift of the fundamental frequency and the attenuation of harmonics. By establishing a voice fatigue model, use the formula Calculate the voice fatigue index. represents the change amount of the fundamental frequency. and respectively represent the change amounts of the amplitude and phase of the w-th harmonic. represents the amplitude of the w-th harmonic at the initial moment. , and are model coefficients, and x is the total number of harmonics.
[0044] Step 4: Multidimensional correlation analysis: By introducing the combination of machine learning algorithms and time series analysis, comprehensively analyze the correlation between voice data and neck and laryngeal muscle movement data.
[0045] The machine learning algorithm sets the voice feature vector as , the neck and laryngeal muscle movement feature vector as , and the prediction function obtained through training as , represents the weight corresponding to the voice feature, represents the weight corresponding to the neck and laryngeal muscle movement feature, b is the bias term, n1 represents the number of elements in the voice feature vector, s i1 represents the i1-th feature in the voice feature vector, k1 represents the number of elements in the neck and laryngeal muscle movement feature vector, m j1 represents the j1-th feature in the neck and laryngeal muscle movement feature vector.
[0046] The machine learning algorithm is used to predict the changes in voice metrics such as the trill rate and amplitude ratio of the voice under specific muscle movement states.
[0047] The time series analysis establishes an ARIMA(p, d, q) model based on the voice fatigue index F t and any parameter of the neck and laryngeal muscle movement. Let F t be after d times of differencing, and establish the model , represents the autoregressive coefficient, represents the moving average, represents the white noise term, p1 represents the autoregressive order, i2 is used to iterate through each lag order from 1 to p1 in the autoregressive part, q1 represents the moving average order, j2 is used to iterate through each lag order from 1 to q1 in the moving average part, and t1 represents the current time.
[0048] The time series analysis analyzes the mutual relationship between the two in the time dimension through the ARIMA(p, d, q) model, and predicts the possible development of the voice fatigue in the next few time points with the continuous change of the average tension of the laryngeal muscles.
[0049] Step 5: Visualization: including a real-time data dashboard and a comparison chart of historical data and predicted trends, which are used to display the data in an intuitive form.
[0050] The real-time data dashboard displays the metrics output in Steps 3 and 4 in a visual way. Each metric is represented by a circular dial, and the pointer points to the current numerical position in real time. Different color regions are divided on the dial according to the set threshold range. Green represents the normal range, yellow represents approaching the warning range, and red represents exceeding the warning range.
[0051] The historical data and predicted trend comparison chart shows the historical data of the singer and the predicted trends for the future by a multi-dimensional correlation model through a comprehensive chart. The horizontal axis of the chart represents time, and the vertical axis represents the values of various indicators. Different colored lines are used to represent the changing trends of historical data and the predicted future trends of the model respectively. The blue line represents the actual data changes of the voice resonance effect indicator in the past month, and the red line represents the development trend of this indicator in the next two weeks based on the model prediction. Annotations and notes are added to the chart to facilitate the user's understanding of the data changes. For possible abnormal situations in the predicted trends, early warning annotations are made in advance. When it is predicted that the voice fatigue index will exceed the warning threshold in the next week, a yellow triangle is marked on the chart, and a prompt is given: "Please note that the voice fatigue may reach a dangerous level in the next week. Please adjust the training plan." And detailed analysis is carried out in a graphical way. Specifically, through a tree chart, the contribution degree of each factor in the voice characteristics and neck and laryngeal muscle movement parameters to the abnormal indicator is shown, and different colors and line thicknesses are used to represent the influence degree. Clicking on each factor in the tree chart can also display detailed historical data.
[0052] The present invention realizes the accurate acquisition of sound signals and physiological data by setting a multi-directional microphone array and neck and laryngeal muscle fiber optic sensors. The microphone array can capture the singer's voice omni-directionally, while the fiber optic sensors can accurately sense the deformation and tension of the neck and laryngeal muscles, providing a high-quality data basis for subsequent analysis. During the data acquisition stage, the system continuously records sound and physiological data at a high sampling rate and marks segments according to different stages of singing. Meanwhile, metadata such as the song name and singer information are also recorded to facilitate the classification management and in-depth analysis of subsequent data. This step ensures the comprehensiveness and accuracy of the data. The analysis of sound data is the core link of this method. By performing multi-dimensional analysis on the collected sound data, including aspects such as dynamic range, timbre change, resonance effect, vocal efficiency, vocal techniques, and adaptability, the system can accurately evaluate the singer's vocal level and existing problems, providing a scientific basis for subsequent personalized teaching and training. By introducing machine learning algorithms and time series analysis, the correlation between sound characteristics and neck and laryngeal muscle movements is deeply explored. This can not only reveal the mystery of the vocalization mechanism but also predict the possible future change trends of sound indicators. Based on these analysis results, teachers can formulate personalized training plans for singers to improve their vocal techniques and performance levels. The analysis results are presented in an intuitive and easy-to-understand manner. Through a real-time data dashboard and a comparison chart of historical data and predicted trends, singers and teachers can clearly see the values and change trends of various indicators. Meanwhile, the system can also provide warnings and detailed analysis windows to help singers discover problems in a timely manner and adjust their training plans. This visual presentation method makes the analysis results more intuitive and easy to understand, facilitating their application in teaching and training practices.
[0053] Secondly: In the accompanying drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments are involved. Other structures can refer to the usual designs. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other;
[0054] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A sound visualization analysis method for vocal music teaching, characterized in that: include: Step 1: Sensor setup: including microphone array and fiber optic sensors of neck and laryngeal muscles for sound collection and physiological data collection; Step 2: Data collection: used to collect sound data and physiological data; Step 3: Sound data analysis: used to analyze the sound data collected in step 2 to obtain dynamic range and timbre change data, resonance and vocal efficiency data, and skill and adaptability data; Step 4: Multidimensional correlation analysis: By introducing a machine learning algorithm and combining it with time series analysis, the correlation between the sound data and the neck and throat muscle movement data is comprehensively analyzed; Step 5: Visualization: This includes real-time data dashboards and historical data and forecast trend comparison charts to convert data into an intuitive form for presentation.
2. The sound visualization analysis method for vocal music teaching according to claim 1, characterized in that: The sound collection uses a multi-directional microphone array, which is arranged in a semicircular track around the singer, about 0.8-1.2 meters away from the singer, and is equipped with a professional-grade audio interface to convert the sound signal into a digital signal and transmit it to the computer; the physiological data collection uses a fiber optic sensor wrapped around the muscle groups of the neck and throat to sense muscle deformation and tension changes; The sound data is collected by the microphone array at a sampling rate of 48kHz and 24-bit quantization accuracy when the singer is practicing vocal music or singing different pieces of music. During the collection process, segmented marking is performed according to different stages of singing for subsequent targeted analysis; the physiological data is recorded by the neck and throat muscle fiber optic sensors at a sampling rate of 1000Hz to record muscle deformation and tension data; The dynamic range and timbre change data include the volume attenuation ratio from the lowest note to the highest note, the timbre brightness index, the timbre consistency when spanning the range, and the vibrato rate and amplitude ratio; the resonance and vocal efficiency data include the head cavity resonance ratio, the degree of vocal cord closure during vocalization, the breath flow and pitch stability coefficient, and the flexibility of resonance cavity conversion; the skill and adaptability data include the accuracy of ornamentation, the rhythm flexibility index, the sound intensity control accuracy, and the sound fatigue index.
3. The sound visualization analysis method for vocal music teaching according to claim 2, characterized in that: The specific method for analyzing the volume attenuation ratio from the lowest sound to the highest sound is: perform time domain analysis on the sound signal, identify the starting and ending positions of the lowest sound and the highest sound, calculate the root mean square value of the volume between the two, and use the formula The volume attenuation ratio is obtained, where is the RMS value of the lowest volume, is the root mean square value of the volume of the highest sound. The specific analysis method of the timbre brightness index is as follows: the sound signal is converted into the frequency domain using Fourier transform, and weights are assigned to different frequency bands according to the human ear's perception characteristics of sounds of different frequencies. Calculate the timbre brightness index, where is the weight of the kth frequency band, is the energy proportion of the frequency band, m represents the number of divided frequency bands, and the specific method for analyzing the consistency of timbre when the range spans is as follows: at the key transition point of the range span, the Mel frequency cepstrum coefficient feature vector of the sound is extracted, and the cosine similarity between adjacent feature vectors is calculated using the formula The timbre consistency index is obtained, where and are the i-th MFCC coefficients of two adjacent transition points, n is the dimension of the MFCC coefficient, and the specific analysis method of the vibrato rate and amplitude ratio is as follows: the instantaneous frequency of the sound signal is extracted by Hilbert transform, the vibrato part is detected, the number of frequency fluctuations per unit time is calculated to obtain the vibrato rate, and the ratio of the maximum amplitude of the frequency fluctuation to the average frequency is taken as the vibrato amplitude ratio, that is, , , Indicates the vibrato rate, Indicates the number of frequency fluctuations, Indicates time, represents the amplitude ratio, Indicates the maximum frequency fluctuation amplitude, Represents the average frequency.
4. The sound visualization analysis method for vocal music teaching according to claim 2, characterized in that: The head cavity resonance ratio analysis method is specifically as follows: based on the resonance peak analysis of the sound signal, the frequency and intensity of the resonance peak are determined by using the linear prediction coding technology, and the energy of the head cavity resonance resonance peak is compared with the total sound energy, and the formula is used. Calculate the head cavity resonance ratio, Indicates the energy of the resonance peak related to the head cavity resonance, Represents the total energy of the sound signal. The specific analysis method of the vocal cord closure during phonation is as follows: Analyze the harmonic structure of the sound signal, calculate the amplitude ratio and phase relationship between the harmonic and the fundamental wave, and use the empirical formula to estimate the vocal cord closure , and represents the amplitude and phase of the jth harmonic, and represents the amplitude and phase of the fundamental wave, and represents the empirical coefficient, p represents the number of harmonics, and the specific analysis method of the breath flow and pitch stability coefficient is as follows: combining the breath flow data collected by the respiratory sensor and the pitch change of the sound signal, calculate the ratio of the standard deviation of the pitch to the average breath flow, and the formula is: , represents the standard deviation of pitch, The analysis method of resonance cavity conversion flexibility is as follows: by monitoring the rapid change of sound signal energy in different frequency bands and the migration of resonance peaks, the change rate of resonance peak frequency and intensity in adjacent time periods is calculated as an indicator of resonance cavity conversion flexibility. The formula is , and are the resonance peak frequency and intensity at the tth moment, q represents the total number of time periods analyzed, Indicates the flexibility of resonance cavity conversion. represents the resonance peak frequency at time t+1, Represents the intensity at time t+1.
5. The sound visualization analysis method for vocal music teaching according to claim 2, characterized in that: The accuracy analysis method of the ornaments is as follows: the ornaments of the performance are compared with the ornaments in the standard score, and the errors in pitch, duration and start time between the two are calculated using the formula Calculate the accuracy index, , as well as They represent the pitch, duration and start time of the standard ornament. , as well as They represent the pitch, duration and start time of the actual singing grace note. , as well as Respectively represent the maximum values of pitch, duration and start time of all ornament parameters. , as well as They represent the minimum values of pitch, duration and start time of all ornamentation parameters respectively, s represents the number of ornamentations, and the analysis method of rhythm flexibility index is as follows: the actual rhythm of the performance is determined by using the beat tracking algorithm, and compared with the standard rhythm, and the statistics of rhythm deviation within a certain time window are calculated. Calculate the rhythm flexibility index, The variance of the rhythm deviation is expressed in the following way: the analysis method of the sound intensity control accuracy is as follows: the dynamic range of the sound signal is analyzed and divided into multiple levels, and the accuracy and delicacy of the conversion of the intensity between different levels in the actual singing is calculated. calculate, Indicates whether the u-th intensity conversion is accurate, 1 for accurate and 0 for inaccurate, v represents the total number of intensity conversions. The specific analysis method of the sound fatigue index is: monitor the long-term change trend of the sound signal, including the drift of the fundamental frequency and the attenuation of the harmonics, and establish a sound fatigue model using the formula Calculate the voice fatigue index, represents the change in fundamental frequency, and Respectively represent the changes in the w-th harmonic amplitude and phase, represents the amplitude of the wth harmonic at the initial moment, , as well as is the model coefficient, and x is the total number of harmonics.
6. The sound visualization analysis method for vocal music teaching according to claim 1, characterized in that: The machine learning algorithm assumes that the sound feature vector is , the neck and laryngeal muscle movement feature vector is , the prediction function obtained through training is , represents the weight corresponding to the sound feature, represents the weight corresponding to the neck and throat muscle movement characteristics, b is the bias term, n1 represents the number of elements in the sound feature vector, s i1 represents the i1th feature in the sound feature vector, k1 represents the number of elements in the neck and throat muscle movement feature vector, m j1 Represents the j1th feature in the neck and laryngeal muscle movement feature vector.
7. The sound visualization analysis method for vocal music teaching according to claim 1, characterized in that: The time series analysis was based on the sound fatigue index F t and any parameter of the neck and throat muscle movement to establish an ARIMA (p, d, q) model. Let F t After d differences, , build the model , represents the free regression coefficient, represents the moving average, represents the white noise term, p1 represents the autoregressive order, i2 is used to traverse each lag order from 1 to p1 in the autoregressive part, q1 represents the moving average order, j2 is used to traverse each lag order from 1 to q1 in the moving average part, and t1 represents the current moment.
8. The sound visualization analysis method for vocal music teaching according to claim 1, characterized in that: The real-time data dashboard displays the indicators output from step 3 and step 4 in a visual manner. Each indicator is represented by a circular dial, and the pointer points to the current numerical position in real time. The dial is divided into different color areas according to the set threshold range. Green indicates the normal range, yellow indicates approaching the warning range, and red indicates exceeding the warning range.
9. The sound visualization analysis method for vocal music teaching according to claim 1, characterized in that: The historical data and predicted trend comparison chart draws a comprehensive chart to show the singer's historical data and the predicted trend of the future by the multi-dimensional correlation model. The horizontal axis of the chart represents time, and the vertical axis represents the values of various indicators. Lines of different colors are used to represent the changing trends of historical data and the future trends predicted by the model. The blue line represents the actual data changes of the sound resonance effect indicator in the past month, and the red line represents the development trend of the indicator in the next two weeks based on the model prediction. Labels and annotations are added to the chart to facilitate users to understand the changes in the data. For abnormal situations that may occur in the predicted trend, early warning annotations are made in advance. When it is predicted that the voice fatigue index will exceed the warning threshold in the next week, it is marked with a yellow triangle on the chart, and the prompt "Please note that the voice fatigue may reach a dangerous level in the next week, please adjust the training plan" is used. A detailed analysis is carried out in a graphical way, specifically, a tree diagram is used to show the contribution of each factor in the sound characteristics and neck and throat muscle movement parameters to the abnormal indicators, and different colors and line thicknesses are used to indicate the degree of influence. Clicking each factor in the tree diagram can also display detailed historical data.
Citation Information
Patent Citations
Singing skill detection system for vocal music teaching
CN106097828A
Vocal singing training method and device
CN114664324A