Beauty sound training intelligent system and training method based on resonance adjustment and respiration feedback

By using an intelligent system based on resonance adjustment and breathing feedback, the system monitors and analyzes breathing and vocal characteristics in vocal training in real time, providing personalized feedback. This solves the problems of delayed feedback and visualization in traditional vocal teaching and improves the efficiency of remote teaching.

CN120998094APending Publication Date: 2025-11-21HUBEI ENG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511135610.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional vocal music teaching lacks systematic technical support and quantifiable analysis methods. Training feedback is delayed and subjective, resonance areas are difficult to visualize, remote teaching is inefficient, physiological parameter monitoring is weak, and real-time supervision and personalized guidance cannot be achieved.

Method used

The system employs an intelligent system based on resonance adjustment and respiratory feedback, which includes a respiratory data acquisition module, a sound acquisition and analysis module, a resonance recognition and projection module, an intelligent training feedback module, a visual interactive client, and a remote teaching platform. It monitors and analyzes respiratory and sound characteristics in real time through wearable sensors and high-sensitivity microphones, and provides personalized feedback by combining deep learning models and acoustic resonance imaging algorithms.

Benefits of technology

It enables scientific monitoring and real-time evaluation of the training process, provides personalized guidance, improves training effectiveness and efficiency, solves the problems of delayed feedback and ambiguous information, and ensures professional guidance for remote teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention relates to a beautiful sound training intelligent system based on resonance adjustment and respiration feedback. The beautiful sound training intelligent system comprises a respiration data acquisition module, a sound acquisition and analysis module, a resonance identification and projection module, an intelligent training feedback module, a visual interaction client, a remote teaching and teaching management platform and a training master control system. The invention further relates to a training method using the beauty training intelligent system based on resonance adjustment and respiration feedback. According to the invention, scientific monitoring, real-time evaluation and personalized guidance of the whole vocal music training process are realized; the problems of feedback lag, subjective judgment, fuzzy information and the like of traditional training are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent vocal music education, in particular to a bel canto training intelligent system and method based on resonance adjustment and breathing feedback. BACKGROUND

[0002] With the development of vocal music education, especially in the teaching of bel canto, breathing control, resonance adjustment and tone stability are considered as the three core links of training. At present, the traditional teaching method is still mainly face-to-face guidance of teachers and imitation learning of students. This method depends on the experience judgment of teachers and the perceptual ability of students, and lacks systematic technical assistance and quantifiable analysis means.

[0003] Specifically, the current teaching has the following main problems: (1) The training feedback is lagging and subjective. Teachers cannot make accurate feedback immediately after the student sings, and students cannot judge the correctness of the current singing; (2) Lack of quantitative record of training process. Key data such as breathing pattern and sound wave characteristics cannot be effectively collected and analyzed, making it difficult to review and track the teaching content; (3) Lack of intuitive guidance for resonance and tone control. The resonance area (such as chest cavity, pharyngeal cavity, nasal cavity, etc.) is abstract for beginners and difficult to locate without visualization; (4) Low efficiency of remote teaching. It is impossible to supervise the training quality and effect evaluation in real time through remote means, resulting in the lack of professional guidance for online vocal music courses; (5) Weak physiological parameter monitoring. Most of the current teaching lacks systematic measurement of physiological factors such as breathing depth and rhythm stability.

[0004] Although there are some auxiliary tools in the market, such as metronome, simple audio analysis software or sports bracelet, these tools are mostly general-purpose devices and cannot meet the complex needs of dynamic coordination and regulation of "breathing-resonance-tone" in professional bel canto training.

[0005] Therefore, there is an urgent need for an intelligent vocal training system that integrates data collection, intelligent analysis, real-time feedback and remote interaction to improve training effectiveness and solve many problems existing in the current teaching. SUMMARY

[0006] The present application provides a bel canto training intelligent system and method based on resonance adjustment and breathing feedback to overcome the problems of feedback lag, subjective judgment and information ambiguity in traditional training.

[0007] To solve the above problems, the technical scheme provided by the present application is: A kind of intelligent system of bel canto training based on resonance adjustment and breathing feedback, including breathing data acquisition module, sound acquisition and analysis module, resonance identification and projection module, intelligent training feedback module, visual interactive client, remote teaching and teaching management platform, training master system, wherein: The breathing data acquisition module monitors the depth of breath, breathing frequency, breathing rhythm, duration of training object in the process of singing in real time by using wearable sensing device, and distinguishes chest breathing and abdominal breathing;The sound acquisition and analysis module collects audio data of the sound of training object by using sound acquisition equipment, and then analyzes the corresponding acoustic characteristics by using built-in audio analysis algorithm, and assists in judging sound region attribution and tone stability;The resonance identification and projection module maps the sound wave resonance area into the three-dimensional head model by using sound wave resonance imaging algorithm, and is used to display the current sound wave concentration area, and assists the training object to adjust the resonance position autonomously;The intelligent training feedback module generates improvement suggestions for the current training object and presents in real time by fusing deep learning model and artificial preset evaluation standard, based on historical training data and the current sound state and breathing matching degree collected by the intelligent system;The visual interactive client is used to display the training data image of the corresponding training object through the smart device of user, and interacts with the training database of the intelligent system through network communication;The remote teaching and teaching management platform is used to assist teachers in remote teaching, manage the training data of training object and interact with training object;The training master system is used to store and manage the training data of the intelligent system.

[0008] Preferably, the wearable sensing device includes chest-abdominal dual-zone breathing sensing belt and skull vibration-guided resonance sensor;The chest-abdominal dual-zone breathing sensing belt includes a belt and a multi-point pressure / displacement sensor;The belt is made of high-elastic flexible material;A plurality of multi-point pressure / displacement sensors are built-in the belt;The multi-point pressure / displacement sensor is used to monitor the breathing movement of the chest and abdomen of the training object in real time;The skull vibration-guided resonance sensor is used to collect skull micro-vibration signals and assist in positioning the distribution area of sound wave resonance point.

[0009] Preferably, the sound acquisition equipment includes array high-sensitivity microphone;The array high-sensitivity microphone is electrically connected with the intelligent system, and is used to transmit the collected sound audio data to the training master system, while controlling the frequency response in the range of 20Hz~20kHz.

[0010] Preferably, the audio analysis algorithm is a spectral analysis by fast Fourier transform and a time-frequency multi-resolution analysis based on wavelet analysis method, and then the acoustic features are extracted and used to assist in judging sound region attribution and tone quality stability; the acoustic features include pitch, formant position, and spectral energy distribution; the audio analysis algorithm includes the following steps: Sa100. The original audio signal is pre-emphasized, expressed as: Wherein: is used to represent the original discrete-time audio signal; is used to represent the pre-emphasized audio signal; is used to represent the pre-emphasis processing coefficient; Sa200. The pre-emphasized audio signal is divided into frames according to a predetermined length and Hamming windowed, specifically expressed as: Wherein: is used to represent the m-th frame windowed signal sequence; is used to represent the frame index; is used to represent the frame shift (frame interval); is used to represent the sample point index in the current frame; is used to represent the length of each frame (window length); Sa300. The result obtained by Sa200 is subjected to fast Fourier transform spectral analysis, and then the amplitude spectrum and power spectrum are obtained, and the frequency corresponding to the maximum power spectrum in the human voice pitch range is found , and converted into frequency; the fast Fourier transform spectral analysis performs discrete Fourier transform on each frame of signal, expressed as: Wherein: is used to represent the frequency index; is used to represent the complex spectrum value (Fourier transform result) of the m-th frame signal at the k-th frequency point; is used to represent the number of points of fast Fourier transform (FFT); The amplitude spectrum is expressed as: The power spectrum is expressed as: The frequency conversion is expressed as: Wherein: a pitch frequency of the m-th frame; a frequency index value corresponding to a dominant frequency (pitch); a sampling frequency; Sa400. Time-frequency multi-resolution analysis is performed based on continuous wavelet transform, and an energy spectrum is obtained through discrete multi-scale decomposition; then local peaks are identified in the energy spectrum, and a formant is located in combination with the result of the fast Fourier transform spectrum analysis in step Sa300; The continuous wavelet transform is expressed by the following formula: Wherein: a scale factor in the wavelet transform; a translation factor in the wavelet transform; a continuous time variable; a wavelet base function; an original audio signal under continuous time; a wavelet transform coefficient; The energy spectrum is expressed by the following formula: Wherein: an energy distribution on a scale ; Sa500. The power spectrum is integrated in a preset frequency band range, and then a frequency spectrum energy distribution feature is obtained, which is expressed by the following formula: Wherein: a power spectrum density value of the m-th frame; a selected frequency band index set; an energy sum of a certain frequency band of the m-th frame.

[0011] Preferably, the acoustic resonance imaging algorithm specifically comprises: acquiring a sound source direction based on a time difference estimation of the array high-sensitivity microphone; constructing a three-dimensional resonance map in combination with a power spectrum density estimation, and mapping to a digital human head model for visual display; and specifically comprising the following steps: Sb100. Pre-emphasizing, framing and windowing each channel signal of the array, and then obtaining a frequency spectrum value corresponding to the microphone channel; Sb200. Performing GCC-PHAT time difference estimation, which is expressed by the following formula: Wherein: The weighted cross-correlation function (PHAT method) used to characterize the i-th microphone and the j-th microphone; Used to characterize the time difference between the sound wave traveling from the sound source to different microphones; Used to characterize the estimated time difference between the i-th microphone and the j-th microphone in the m-th frame; Sb300. The direction of the sound source is obtained by solving the equation, and is expressed as follows: in: Used to characterize the elevation angle of the incident sound wave relative to the array coordinate axes; Used to characterize the azimuth angle of the incident sound wave relative to the array coordinate axes; Used to characterize the direction of a sound source; The position vector used to characterize the microphone; Used to characterize the speed of sound; Used to characterize the estimated time difference between the i-th microphone and the j-th microphone; Sb400 is estimated by combining the power spectral density, and expressed as follows: in: Used to characterize the The temporal sampling signal of each channel in the m-th frame; Used to characterize the normalization factor; Used to characterize the power spectral density estimate of the m-th frame; Sb500. Construct the three-dimensional resonance diagram; repeatedly solve for the sound source direction and estimate the power spectral density for each direction of the voxel mesh, and then obtain the energy estimate of the sound source point, expressed by the following formula: in: Used to characterize the sound source point Energy estimates; Sb600. Map the three-dimensional resonance map constructed in step Sb00 onto the digital head model and visualize it, specifically expressed as follows: in: Used to characterize the first in three-dimensional space The coordinate vector of an individual element; Used to characterize the Three-dimensional spatial coordinate vectors of candidate sound source points; The scaling parameter used to characterize the Gaussian kernel function; Used to characterize voxels The resonance intensity value.

[0012] Preferably, the deep learning model is specifically a multi-modal feature fusion model that fuses convolutional neural network to extract acoustic features and recurrent neural network to capture timing dynamics; the deep learning model gives intelligent scores and suggestions for different sound regions and breathing patterns, specifically including the following steps: Sc100. Constructing an input layer of the deep learning model, expressed as follows: Wherein: A is used to represent the acoustic feature sequence; B is used to represent the breathing feature sequence; Sc200. Layer normalization and position encoding are performed on the input layer, expressed as follows: Wherein: is used to represent the feature corresponding to the energy envelope of the resonance region; is used to represent the feature corresponding to the energy envelope of the breath region; Sc300. Constructing an acoustic branch, expressed as follows: Wherein: is used to represent the feature vector at time step t after the first layer convolution processing; is used to represent the first layer convolution kernel parameter; is used to represent the number of stacked layers of the convolution layer in the model; is used to represent the linear rectifier activation function; Sc400. Constructing a breathing branch, expressed as follows: Wherein: GRU is used to represent the gating recurrent unit operation; is used to represent the forward GRU output; is used to represent the reverse GRU output; is used to represent the spliced output of the bidirectional GRU; Sc500. Calculate the cross-modal attention, expressed as follows: Wherein: is used to represent the query matrix of the resonance channel; is used to represent the key matrix of the breathing channel; is used to represent the vector dimension; is used to represent the value matrix of the breathing channel; for characterizing the attention normalization function; for characterizing the resonance-breathing interaction result calculated by the h-th attention channel; Sc600. The calculated to feedforward network with residual, expressed as follows: wherein: for characterizing the output feature of the multi-head attention module as the reference baseline input of the feedforward module; for characterizing the output of the feedforward module; for characterizing the final fusion feature at time step ; for characterizing the first layer bias term; for characterizing the second layer bias term; ReLU for characterizing the activation function; LayerNorm for characterizing the layer normalization; for characterizing the first layer fully connected weight; for characterizing the second layer fully connected weight; Sc700. The intelligent scoring and suggestion output for different sound zones and breathing patterns, expressed as follows: , wherein: for characterizing the final feature vector fused from the resonance and breathing channels; for characterizing the weight matrix; for characterizing the offset for adjusting the scoring output; for characterizing the comprehensive score; for characterizing the attention mapping weight matrix; for characterizing the attention weight of the k-th suggestion; for characterizing the k-th suggestion vector; for characterizing the final suggestion vector; Sc800. The loss function of the deep learning model is calculated, specifically expressed as follows: wherein: for characterizing the scoring loss weight coefficient; for characterizing the scoring loss function; for characterizing the suggestion loss weight coefficient; for characterizing the suggestion branch loss function; for characterizing the regularization term weight coefficient; for characterizing the model parameter vector; for characterizing the total loss function.

[0013] Preferably, the artificial preset evaluation criteria are used for the auxiliary criteria of system output, including resonance peak deviation, timbre brightness, breath support degree and comprehensive score. The resonance peak deviation is expressed by the following formula: Wherein: for representing the nth resonance peak frequency (measured value); for representing the nth resonance peak frequency (reference value); The timbre brightness is expressed by the following formula: Wherein: for representing the kth frequency point frequency; for representing the kth frequency point power spectrum value; The breath support degree is expressed by the following formula: Wherein: for representing the breath maintenance ratio; for representing the breath pressure stability score; The comprehensive score is expressed by the following formula: Wherein: for representing the resonance weight; for representing the timbre weight; for representing the breath weight; for representing the maximum allowable resonance deviation; for representing the maximum frequency.

[0014] A training method using the intelligent system for bel canto training based on resonance adjustment and breath feedback, comprising the following steps: S100. Collect the basic information of the current training object, and initialize the settings of the intelligent system; S200. Wear the wearable sensing device for the current training object; at the same time, set the sound collecting equipment within the preset range of the current user's location; then connect the wearable sensing device and the sound collecting equipment with the training master control system; S300. Real-time read the respiratory fluctuation signal of chest and abdomen through the breath data acquisition module, and automatically adjust the sampling frequency to adapt to different teaching rhythm; the respiratory fluctuation signal includes the breath depth, the breath frequency, the breath rhythm and the duration; S400. Perform the filtering and normalization processing on the respiratory fluctuation signal collected in step S300; S500. Collect the audio signal of the current training object through the sound collection and analysis module, and then extract the acoustic features through Fourier transform and wave transform; S600. Upload the respiratory and sound feature data obtained after processing by the respiratory data collection module and the sound collection and analysis module to the training master system for fusion analysis, and synchronously return to the visualization interactive client display in real time; S700. The intelligent training feedback module built-in deep learning model and the artificial preset evaluation standard are used to intelligently evaluate the training data processed by the training master system in multiple dimensions, and the evaluation results of respiratory-voice synchronization degree, resonance concentration degree, pitch control accuracy and breath support matching degree are generated. Then generate and feedback to the visualization interactive client and the remote teaching and teaching management platform in the form of waveform graph, heat map and stability curve; S800. Teachers intelligently manage the training process and training data of the current training object through the remote teaching and teaching management platform, which specifically includes real-time remote viewing of user data and training data of the training object, uploading demonstration content, issuing tasks for the current training object. The training master system automatically generates training progress reports and growth curves to provide stage feedback and next training suggestions for the training object.

[0015] Preferably, the filtering and normalization processing in step S400 specifically includes the following steps: Sd100. Use a band-pass filter to filter out non-respiratory bandwidth; Sd200. Apply Kalman filter to smooth the sudden artifacts; Sd300. Z-score normalization is performed on the filtered signal, so as to eliminate individual differences and ensure that the data of different training objects are comparable.

[0016] Preferably, the multi-dimensional evaluation results include the synchronization degree of breathing and voice, whether the resonance area is concentrated and stable, the accuracy and stability of pitch control, whether the breath support matches the length and intensity of the voice, the breath matching degree score, and the training state personalized prompt card. The output results of the multi-dimensional evaluation results after image processing include the respiratory waveform graph, the acoustic spectrum heat map, the resonance distribution three-dimensional graph and the pitch stability curve.

[0017] Compared with the prior art, the present application has the following advantages: Since the present application integrates physiological and acoustic dual-dimensional data analysis, it realizes the quantification and objectivity evaluation of training effect, thereby solving the problem of lack of objective and quantitative evaluation method in the training feedback of the prior art.

[0018] The application dynamically adjusts the training strategy according to the vocal characteristics and historical progress of the trainee, thereby realizing personalized training path planning and improving training efficiency.

[0019] The application automatically prompts the problems of vocal deviation and abnormal breathing during the training process, thereby improving the efficiency and solving the problem of lagging training feedback in the prior art.

[0020] The application makes intuitive feedback on the training process of the training object through visual graphics, introduces a three-dimensional resonance surface visualization technology based on power spectrum density (PSD) and sound source direction estimation, and combines a digital head model, thereby solving the problem of abstract resonance area and the inability to provide visual feedback during training in the prior art.

[0021] The application adapts to the new mode of network teaching, ensures the controllable synchronization of teaching quality in different places, and thereby solves the problem of the lack of professional guidance in online vocal music courses due to the inability to supervise the training quality and effect evaluation in real time through remote means in the prior art.

[0022] The application adds a high-sensitivity respiratory airflow acquisition module based on the existing single-channel acoustic acquisition, thereby realizing the synchronous and high-precision acquisition of acoustic and respiratory data.

[0023] The application combines convolutional neural network (CNN) and recurrent neural network (RNN) structures to jointly analyze acoustic features and respiratory features, thereby realizing multi-dimensional analysis of resonance distribution, pitch stability, and breath support.

[0024] The application realizes targeted analysis and personalized suggestion output for different voice parts, different voice regions, and different breathing patterns through an expert knowledge base and artificial preset evaluation standards, thereby solving the problem of the lack of personalized guidance in traditional methods. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 The figure is a structure diagram of the bel canto training intelligent system based on resonance adjustment and breathing feedback of the specific embodiment of the application. Figure 2 The figure is a schematic diagram of a wearable sensing device of the specific embodiment of the application. Figure 3 The figure is a three-dimensional analysis interface diagram of the system resonance distribution of the specific embodiment of the application. Figure 4 The figure is a breathing-vocalization synchronization and stability analysis interface diagram of the specific embodiment of the application.

[0026] Wherein: 100. respiratory data acquisition module, 200. sound acquisition and analysis module, 300. resonance recognition and projection module, 400. intelligent training feedback module, 500. training master system, 600. visual interactive client, 110. wearable sensing device, 111. dual-zone chest and abdomen respiratory sensing belt, 111.1. belt, 111.2. multi-point pressure / displacement sensor DETAILED DESCRIPTION

[0027] The present application will be further illustrated below in conjunction with specific examples, which are intended to illustrate the present application but not to limit the scope of the present application. After reading the present application, those skilled in the art can make various modifications to the equivalent forms of the present application, which fall within the scope defined by the appended claims.

[0028] The present application claims an intelligent system for bel canto training based on resonance adjustment and respiratory feedback, as shown in Figure 1 The present application claims an intelligent system for bel canto training based on resonance adjustment and respiratory feedback, as shown in The respiratory data acquisition module 100 monitors the respiratory depth, respiratory frequency, respiratory rhythm, duration of the training object during singing in real time by using the wearable sensing device 110, and distinguishes between chest breathing and abdominal breathing, especially suitable for the abdominal support requirement emphasized in bel canto training; the sound acquisition and analysis module 200 collects audio data of the training object's voice by using sound acquisition equipment, then analyzes the corresponding acoustic characteristics through the built-in audio analysis algorithm, and assists in judging the sound region attribution and tone stability; the resonance recognition and projection module 300 maps the sound wave resonance area to the three-dimensional head model by using the sound wave resonance imaging algorithm, and is used to display the current sound wave concentration area, intuitively showing whether the current sound wave is concentrated in the chest cavity, pharyngeal cavity, nasal cavity or head cavity, and assisting the training object to adjust the resonance position independently; the intelligent training feedback module 400 fuses the deep learning model and the artificially preset evaluation standard, and simultaneously generates improvement suggestions for the current training object based on the historical training data and the current voice state and respiratory matching degree collected by the intelligent system, and presents in real time; the visual interactive client 600 is used to display the training data image of the corresponding training object through the user's smart device, and interacts with the training database of the intelligent system through network communication; the remote teaching and teaching management platform is mainly used to assist teachers in remote teaching, manage the training data of the training object, and interact with the training object. Teachers can log in to the management system to view the training records of the training object, play practice audio and video, make real-time comments or leave messages, issue personalized tasks, and build a complete "learning-practice-evaluation-improvement" closed-loop system; the training master system 500 is used to store and manage the training data of the intelligent system.

[0029] It should be noted that, as shown in Figure 2 The wearable sensing device 110 includes a chest-abdominal dual-zone respiration sensing belt 111 and a skull vibration-guided resonance sensor; the chest-abdominal dual-zone respiration sensing belt 111 includes a belt 111.1 and a multi-point pressure / displacement sensor 111.2; the belt 111.1 is made of high-elastic flexible material; a plurality of multi-point pressure / displacement sensors 111.2 are built in the belt 111.1; the multi-point pressure / displacement sensor 111.2 is used to monitor the respiration movement of the chest and abdomen of the training object in real time; the sensing belt is made of high-elastic flexible material, which ensures good adhesion to the skin and does not affect the comfort of respiration; the skull vibration-guided resonance sensor is used to collect the skull micro-vibration signal and assist in positioning the distribution area of the acoustic resonance point.

[0030] It should be noted that the sound collecting device includes an array type high-sensitivity microphone; the array type high-sensitivity microphone is electrically coupled with the intelligent system and is used to transmit the collected audio data of the sound to the training master system 500, while controlling the frequency response in the range of 20Hz~20kHz. The sound collecting device should be installed about 1 meter in front of the student to ensure good signal-to-noise ratio. All devices are wirelessly connected to the training master system 500 through Bluetooth 5.0 or Wi-Fi 6, ensuring high bandwidth and low delay data transmission.

[0031] It should be noted that the audio analysis algorithm is a frequency spectrum analysis through fast Fourier transform and a time-frequency multi-resolution analysis based on wavelet analysis method, and then acoustic features are extracted to assist in judging sound region attribution and tone stability; the acoustic features include pitch, formant position, and spectral energy distribution; the audio analysis algorithm includes the following steps: Sa100. Pre-emphasis processing is performed on the original audio signal, which is expressed by formula (1): (1) Wherein: is used to represent the original discrete-time audio signal; is used to represent the audio signal after pre-emphasis processing; is used to represent the pre-processing coefficient.

[0032] Sa200. The audio signal after pre-emphasis processing is divided into frames according to the preset length, and a Hamming window is added, which is specifically expressed by formula (2): (2) Wherein: is used to represent the m-th frame of windowed signal sequence; is used to represent the frame index; is used to represent the frame shift (frame spacing); used to represent the index of the sampling point within the current frame; used to represent the length of each frame (window length).

[0033] Sa300. Fast Fourier transform spectrum analysis is performed on the result obtained by Sa200, thereby obtaining the amplitude spectrum and the power spectrum, and finding the maximum power spectrum corresponding to the pitch in the high range of human voice and converting it into frequency; Fast Fourier transform spectrum analysis performs discrete Fourier transform on each frame of signal, expressed by formula (3): (3) wherein: used to represent the index of the frequency; used to represent the complex spectrum value (Fourier transform result) of the mth frame of signal at the kth frequency point; used to represent the number of points of the fast Fourier transform (FFT).

[0034] The amplitude spectrum is expressed by formula (4): (4) The power spectrum is expressed by formula (5): (5) The frequency conversion is expressed by formula (6): (6) wherein: used to represent the estimated pitch frequency of the mth frame; used to represent the frequency index value corresponding to the fundamental frequency (pitch); used to represent the sampling frequency.

[0035] Sa400. Time-frequency multi-resolution analysis is performed based on continuous wavelet transform, and the energy spectrum is obtained through discrete multi-scale decomposition; then local peaks are identified in the energy spectrum, and the formant is located in combination with the result of the fast Fourier transform spectrum analysis of step Sa300.

[0036] The continuous wavelet transform is expressed by formula (7): (7) wherein: used to represent the scale factor in the wavelet transform; used to represent the translation factor in the wavelet transform; used to represent the continuous time variable; used to represent the wavelet base function; used to represent the original audio signal under continuous time; used to represent the wavelet transform coefficient.

[0037] The energy spectrum is expressed by formula (8): (8) Wherein: is used to represent the energy distribution on the scale .

[0038] Sa500. The power spectrum is integrated in a preset frequency band range, and the frequency spectrum energy distribution feature is obtained, which is expressed by formula (9): (9) Wherein: is used to represent the power spectrum density value of the mth frame; is used to represent the selected frequency band index set; is used to represent the energy sum of a certain frequency band of the mth frame.

[0039] It should be noted that the acoustic resonance imaging algorithm is specifically: based on the time difference estimation of the array high-sensitivity microphone to obtain the sound source direction; combined with the power spectrum density estimation to construct a three-dimensional resonance map, and mapped to the digital human head model for visual display; specifically including the following steps: Sb100. The signals of each channel of the array are pre-emphasized, framed and windowed, and then the frequency spectrum values corresponding to the microphone channels are obtained.

[0040] Sb200. GCC-PHAT time difference estimation is performed, which is expressed by formula (10) and formula (11): (10) (11) Wherein: is used to represent the weighted cross-correlation function (PHAT method) of the ith microphone and the jth microphone; is used to represent the time difference between the sound waves transmitted from the sound source to different microphones; is used to represent the estimated time difference between the ith microphone and the jth microphone in the mth frame.

[0041] Sb300. The sound source direction is solved and expressed by formula (12): (12) Wherein: is used to represent the elevation angle of the incident sound wave and the array coordinate axis; is used to represent the azimuth angle of the incident sound wave and the array coordinate axis; is used to represent the sound source direction; is used to represent the position vector of the microphone; is used to represent the sound velocity; Used to characterize the estimated time difference between the i-th microphone and the j-th microphone.

[0042] Sb400. Estimated using power spectral density, expressed as in equation (13): (13) in: Used to characterize the The temporal sampling signal of each channel in the m-th frame; Used to characterize the normalization factor; Used to characterize the power spectral density estimate of the m-th frame.

[0043] Sb500. Construct a three-dimensional resonance diagram; repeatedly solve for the sound source direction and estimate the power spectral density for each direction of the voxel mesh, and then obtain the energy estimate of the sound source point as expressed by equation (14): (14) in: Used to characterize the sound source point Energy estimates.

[0044] Sb600. Map the three-dimensional resonance diagram constructed in step Sb00 onto the digital head model and visualize it, specifically expressed as in equation (15): (15) in: Used to characterize the first in three-dimensional space The coordinate vector of an individual element; Used to characterize the Three-dimensional spatial coordinate vectors of candidate sound source points; The scaling parameter used to characterize the Gaussian kernel function; Used to characterize voxels The resonance intensity value.

[0045] It should be noted that the deep learning model is specifically a multimodal feature fusion model that combines convolutional neural networks to extract acoustic features with recurrent neural networks to capture temporal dynamics. The deep learning model provides intelligent scoring and suggestions for different vocal registers and breathing patterns, specifically including the following steps: Sc100. Construct the input layer of the deep learning model, expressed by equations (16) and (17): (16) (17) Where: A is used to characterize acoustic feature sequences; B is used to characterize respiratory feature sequences.

[0046] Sc200. Layer normalization and position encoding are performed on the input layer, expressed as formula (18) and formula (19): (18) (19) Wherein: for representing the energy envelope corresponding to the resonance region; for representing the energy envelope corresponding to the breath region. for representing the energy envelope corresponding to the resonance region; for representing the energy envelope corresponding to the breath region.

[0047] Sc300. Build the acoustic branch, expressed as formula (20): (20) Wherein: for representing the feature vector of time step t after the first layer convolution processing; for representing the first layer convolution kernel parameter; for representing the number of stacked layers of the convolution layer in the model; for representing the linear rectification activation function.

[0048] Sc400. Build the breathing branch, expressed as formula (21): (21) Wherein: GRU is used to represent the operation of the gated recurrent unit; for representing the forward GRU output; for representing the reverse GRU output; for representing the spliced output of the bidirectional GRU.

[0049] Sc500. Calculate the cross-modal attention, expressed as formula (22): (22) Wherein: for representing the query matrix of the resonance channel; for representing the key matrix of the breathing channel; for representing the vector dimension; for representing the value matrix of the breathing channel; for representing the attention normalization function; for representing the resonance-breath interaction result calculated by the hth attention channel.

[0050] Sc600. Calculate the to feedforward network and residual, expressed as formula (23): (23) Wherein: for characterizing the output features of the multi-head attention module as a reference baseline for the input of the feed-forward module; for characterizing the output of the feed-forward module; for characterizing the final fusion feature at the time step for characterizing the bias term of the first layer; for characterizing the bias term of the second layer; ReLU for characterizing the activation function; LayerNorm for characterizing the layer normalization; for characterizing the fully connected weight of the first layer; for characterizing the fully connected weight of the second layer.

[0051] Sc700. The intelligent scoring and suggestion are output for different acoustic zones and breathing patterns, expressed as formula (24) and formula (25): , (24) (25) Wherein: for characterizing the final feature vector of the fusion of the resonance and breathing channels; for characterizing the weight matrix; for characterizing the offset of the adjusted score output; for characterizing the comprehensive score; for characterizing the attention mapping weight matrix; for characterizing the attention weight of the kth suggestion; for characterizing the kth suggestion vector; for characterizing the final suggestion vector.

[0052] Sc800. The loss function of the deep learning model is calculated, specifically expressed as formula (26): (26) Wherein: for characterizing the scoring loss weight coefficient; for characterizing the scoring loss function; for characterizing the suggestion loss weight coefficient; for characterizing the suggestion branch loss function; for characterizing the regularization term weight coefficient; for characterizing the model parameter vector; for characterizing the total loss function.

[0053] It should be noted that the artificially preset evaluation standard is used as an auxiliary criterion for system output, including formant deviation, timbre brightness, breath support, and comprehensive score.

[0054] The formant deviation is expressed as formula (27):​ (27) wherein: for representing the n-th formant frequency (measured value); for representing the n-th formant frequency (reference value); the brightness of voice quality is expressed by formula (28): (28) wherein: for representing the k-th frequency point frequency; for representing the k-th frequency point power spectrum value; the breath support degree is expressed by formula (29): (29) wherein: for representing the breath maintenance ratio; for representing the breath pressure stability score; the comprehensive score is expressed by formula (30): (30) wherein: for representing the resonance weight; for representing the voice quality weight; for representing the breath weight; for representing the maximum allowable resonance deviation; for representing the maximum frequency.

[0055] It should be noted that the teacher end can realize the following functions through the remote teaching and teaching management platform: Real-time remote viewing of student data: including historical records, current training conditions, and audio-video synchronous playback.

[0056] Upload demonstration content: such as classic singing sections, technical explanation videos.

[0057] Publish personalized tasks: such as sound area exercises, breath control tasks.

[0058] Automatic scoring and comments: teachers can directly mark problem positions and add voice and text suggestions on the graphical interface.

[0059] Generate training progress reports and growth curves to provide students with stage feedback.

[0060] It should be noted that the present application includes the following application scenarios: Case one (classroom synchronous feedback): in the teaching scene of a music college of a certain university, students wear devices to sing an aria, and the system synchronously displays the breath curve and resonance diagram to assist teachers to adjust teaching strategies in real time. Case two (self-practice mode at home): the student practices independently at home, and the system scores and records his / her pitch, resonance, and breathing condition. The teacher can remotely check and give suggestions after class. Case three (cross-regional remote teaching): a vocal training institution uses the system to guide multiple students distributed in different cities at the same time. The training data and teaching arrangement are managed through the cloud platform to improve teaching efficiency and coverage.

[0061] A training method using a bel canto training intelligent system based on resonance adjustment and breathing feedback, comprising the following steps: S100. Collect the basic information of the current training object and initialize the intelligent system.

[0062] S200. Wear the wearable sensing device 110 for the current training object; at the same time, set the sound collecting equipment within the preset range of the current user's location; then connect the wearable sensing device 110 and the sound collecting equipment with the training master control system 500.

[0063] S300. Real-time read the breathing fluctuation signal of the chest and abdomen through the breathing data acquisition module 100, and automatically adjust the sampling frequency to adapt to different teaching rhythms; the breathing fluctuation signal includes breathing depth, breathing frequency, breathing rhythm, and duration.

[0064] S400. Filter and normalize the breathing fluctuation signal collected in step S300.

[0065] S500. Collect the audio signal of the current training object through the sound collecting and analyzing module 200, and then extract the acoustic features through Fourier transform and wave transform.

[0066] S600. Upload the breathing and sound feature data obtained by processing through the breathing data acquisition module 100 and the sound collecting and analyzing module 200 to the training master control system 500 for fusion analysis, and real-time synchronization back to the visualization interactive client 600 for display.

[0067] S700. Intelligent multi-dimensional evaluation of the training data processed by the training master control system 500 through the deep learning model built-in the intelligent training feedback module 400 and the artificially preset evaluation standard, to generate the evaluation results of breathing-sound synchronization degree, resonance concentration degree, pitch control accuracy, and breath support matching degree; then generate waveform graph, heat map, and stability curve and feedback to the visualization interactive client 600 and the remote teaching and teaching management platform.

[0068] S800. The teacher intelligently manages the training process and training data of the current training object through the remote teaching and teaching management platform, specifically including real-time remote viewing of user data and training data of the training object, uploading demonstration content, and issuing tasks for the current training object; the training progress report and growth curve are automatically generated by the training master control system 500 to provide stage feedback and next training suggestions for the training object.

[0069] It should be noted that the filtering and normalization processing in step S400 specifically includes the following steps: Sd100. Using a band-pass filter to filter out non-breathing bandwidth.

[0070] Sd200. Apply Kalman filter to smooth the sudden artifacts.

[0071] Sd300. Z-score normalization is performed on the filtered signal to eliminate individual differences and ensure comparability of data from different training objects.

[0072] It should be noted that the multi-dimensional evaluation results include the synchronization degree of breathing and sound production, whether the resonance area is concentrated and stable, the accuracy and stability of pitch control, whether the breath support matches the length and intensity of sound production, breath matching degree score, and training state individualized prompt card; the output results after the multi-dimensional evaluation results are visualized include breathing waveform graph, acoustic spectrum thermodynamic map, resonance distribution three-dimensional graph, and pitch stability curve. As shown in the resonance distribution three-dimensional analysis interface Figure 3 , which displays the three-dimensional distribution surface of frequency (Hz), direction (°), and resonance energy (dB), and displays real-time pitch, breath support degree, and other numerical indicators on the right side of the interface. As shown in the breathing-sound production synchronization and stability analysis interface Figure 4 , which displays the time variation of the breathing waveform and pitch deviation curve, and provides real-time evaluation results such as breath support matching degree and resonance concentration degree on the right side of the interface.

[0073] It should be noted that the present application includes the following embodiments: Scheme one: The singer wears a breathing sensor and is located within the acquisition range of the multi-channel microphone array, the system starts multi-channel synchronous acquisition, extracts acoustic characteristics such as pitch, formant, and spectral energy distribution through fast Fourier transform (FFT) and wavelet transform, calculates the sound source direction through time difference estimation (TDOA) and generates a three-dimensional resonance distribution graph combined with power spectral density (PSD); then the acoustic characteristics and breathing characteristics are input into the deep learning analysis module of the fusion CNN and RNN to obtain quantitative scores and training suggestions for resonance adjustment and breath support, and the three-dimensional resonance surface, breathing waveform, and pitch deviation curve are displayed in real time on the interface.

[0074] Scheme two: on the basis of scheme one, load the personalized evaluation standard library, and dynamically adjust the scoring weight and analysis algorithm parameters according to the voice part, song, and singing style of the singer, to realize adaptive training guidance.

[0075] The above schemes are verified by experiments and can significantly improve the ability of the singer in resonance control and breathing support stability.

[0076] In the above detailed description, various features are combined together in a single embodiment to simplify the disclosure. This disclosure method should not be interpreted as reflecting the intention that the embodiments of the claimed subject matter require more features clearly stated in each claim. On the contrary, as reflected in the appended claims, the present application is in a state of less than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby incorporated into the detailed description, where each claim is separately a preferred embodiment of the present application.

[0077] To enable any person skilled in the art to implement or use the present application, the above discloses the disclosed embodiments. For those skilled in the art, various modifications of these embodiments are obvious, and the general principles defined herein can also be applied to other embodiments without departing from the spirit and protection scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in the present application.

[0078] The above description includes an example of one or more embodiments. Of course, it is impossible to describe all possible combinations of components or methods for describing the above embodiments, but those skilled in the art should recognize that various embodiments can be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications and variations falling within the protection scope of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, the coverage of the term is similar to the term "including", as explained in the context of the term "including" used as a conjunction word in the claims. In addition, the use of any one term "or" in the specification of the claims is to mean "non-exclusive or".

[0079] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A smart vocal training system based on resonance adjustment and breathing feedback, characterized in that: It includes a respiratory data acquisition module (100), a sound acquisition and analysis module (200), a resonance recognition and projection module (300), an intelligent training feedback module (400), a visual interactive client (600), a remote teaching and teaching management platform, and a training control system (500), among which: The breathing data acquisition module (100) uses a wearable sensing device (110) to monitor the breathing depth, breathing frequency, breathing rhythm, and duration of the trainee during singing in real time, and distinguishes between chest breathing and abdominal breathing; the sound acquisition and analysis module (200) uses a sound acquisition device to collect audio data of the trainee's vocalization, and then uses a built-in audio analysis algorithm to analyze the corresponding acoustic characteristics and assist in determining the vocal register and sound quality stability; the resonance recognition and projection module (300) uses a sound wave resonance imaging algorithm to map the sound wave resonance area onto a three-dimensional head model, and uses this to display the area where the sound waves are currently concentrated, and assists the trainee in adjusting the resonance position independently; The intelligent training feedback module (400) integrates deep learning models and manually preset evaluation criteria, and generates improvement suggestions for the current training object based on historical training data and the current vocal state and breathing matching degree collected by the intelligent system, and presents them in real time; the visual interactive client (600) is used to display the training data image of the corresponding training object through the user's smart device, and interacts with the training database of the intelligent system through network communication; the remote teaching and teaching management platform is mainly used to assist teachers in remote teaching, manage the training data of training objects, and interact with training objects; the training master control system (500) is used to store and manage the training data of the intelligent system.

2. The intelligent vocal training system based on resonance adjustment and breathing feedback according to claim 1, characterized in that: The wearable sensing device (110) includes a chest and abdomen dual-zone respiratory sensing band (111) and a skull vibration resonance sensor; the chest and abdomen dual-zone respiratory sensing band (111) includes a strap (111.1) and a multi-point pressure / displacement sensor (111.2); the strap (111.1) is made of a highly elastic and flexible material; multiple multi-point pressure / displacement sensors (111.2) are built into the strap (111.1); the multi-point pressure / displacement sensor (111.2) is used to monitor the respiratory movements of the chest and abdomen of the training subject in real time; the skull vibration resonance sensor is used to collect micro-vibration signals of the skull and assist in locating the distribution area of ​​sound wave resonance points.

3. The intelligent vocal training system based on resonance adjustment and breathing feedback according to claim 1, characterized in that: The sound acquisition device includes an array of high-sensitivity microphones; the array of high-sensitivity microphones is electrically coupled to the intelligent system and is used to transmit the collected audio data of the sound to the training main control system (500), while controlling the frequency response within the range of 20Hz to 20kHz.

4. The intelligent vocal training system based on resonance adjustment and breathing feedback according to claim 1, characterized in that: The audio analysis algorithm performs spectral analysis using Fast Fourier Transform and time-frequency multi-resolution analysis based on wavelet analysis to extract acoustic features, which are then used to assist in determining sound zone affiliation and sound quality stability. The acoustic features include pitch, formant positions, and spectral energy distribution. The audio analysis algorithm includes the following steps: Sa100. Pre-emphasis processing is applied to the original audio signal, expressed as follows: in: Used to characterize the original discrete-time audio signal; Used to characterize the audio signal after pre-emphasis processing; Used to characterize the pre-emphasis treatment coefficient; Sa200. The pre-emphasized audio signal is divided into frames according to a preset length, and a Hamming window is added, specifically expressed by the following formula: in: Used to characterize the signal sequence after windowing in the m-th frame; Used to characterize the frame index; Used to characterize frame shift (frame spacing); Used to represent the index of the sampling point within the current frame; Used to represent the length of each frame (window length); Sa300. Fast Fourier Transform (FFT) spectral analysis was performed on the results obtained from Sa200 to obtain the amplitude spectrum and power spectrum, and the maximum power spectrum corresponding to the highest vocal pitch was then located. The frequency is converted to the frequency; the Fast Fourier Transform spectral analysis performs a Discrete Fourier Transform on each frame of signal, expressed by the following formula: in: Used to characterize frequency index; Used to characterize the complex spectrum value (Fourier transform result) of the m-th frame signal at the k-th frequency point; The number of points used to characterize the Fast Fourier Transform (FFT); The amplitude spectrum is expressed by the following formula: The power spectrum is expressed by the following formula: The Frequency conversion is expressed as follows: in: Used to characterize the estimated pitch frequency of the m-th frame; Frequency index values ​​used to characterize the dominant frequency (pitch); Used to characterize the sampling frequency; Sa400. Time-frequency multi-resolution analysis is performed based on continuous wavelet transform, and the energy spectrum is obtained through discrete multi-scale decomposition; then, local peaks are identified in the energy spectrum, and resonance peaks are located by combining the results of the fast Fourier transform spectrum analysis in step Sa300. The continuous wavelet transform is expressed by the following formula: in: Used to characterize the scaling factor in wavelet transform; Used to characterize the translation factor in wavelet transform; Used to characterize continuous-time variables; Used to characterize wavelet basis functions; Used to characterize the raw audio signal over continuous time; Used to characterize wavelet transform coefficients; The energy spectrum is expressed by the following formula: in: Used to characterize scale Energy distribution on; Sa500. Integrate the power spectrum within a preset frequency band to obtain the spectral energy distribution characteristics, expressed by the following formula: in: The power spectral density value used to characterize the m-th frame; Used to characterize the selected set of frequency band indices; Used to characterize the total energy of a certain frequency band in the m-th frame.

5. The intelligent vocal training system based on resonance adjustment and breathing feedback according to claim 1, characterized in that: The acoustic resonance imaging algorithm specifically comprises: obtaining the sound source direction based on the time difference estimation of the array-type high-sensitivity microphone; constructing a three-dimensional resonance map by combining power spectral density estimation, and mapping it to a digital head model for visualization; specifically including the following steps: Sb100. Pre-emphasizes, frames, and windows the signal of each channel of the array to obtain the spectrum value corresponding to the microphone channel; Sb200. Perform GCC-PHAT time difference estimation, expressed as follows: in: The weighted cross-correlation function (PHAT method) used to characterize the i-th microphone and the j-th microphone. Used to characterize the time difference between the sound wave traveling from the sound source to different microphones; Used to characterize the estimated time difference between the i-th microphone and the j-th microphone in the m-th frame; Sb300. The direction of the sound source is obtained by solving the equation, and is expressed as follows: in: Used to characterize the elevation angle of the incident sound wave relative to the array coordinate axes; Used to characterize the azimuth angle of the incident sound wave relative to the array coordinate axes; Used to characterize the direction of a sound source; The position vector used to characterize the microphone; Used to characterize the speed of sound; Used to characterize the estimated time difference between the i-th microphone and the j-th microphone; Sb400 is estimated by combining the power spectral density, and expressed as follows: in: Used to characterize the The temporal sampling signal of each channel in the m-th frame; Used to characterize the normalization factor; Used to characterize the power spectral density estimate of the m-th frame; Sb500. Construct the three-dimensional resonance diagram; repeatedly solve for the sound source direction and estimate the power spectral density for each direction of the voxel mesh, and then obtain the energy estimate of the sound source point, expressed by the following formula: in: Used to characterize the sound source point Energy estimates; Sb600. Map the three-dimensional resonance map constructed in step Sb00 onto the digital head model and visualize it, specifically expressed as follows: in: Used to characterize the first in three-dimensional space The coordinate vector of an individual element; Used to characterize the Three-dimensional spatial coordinate vectors of candidate sound source points; The scaling parameter used to characterize the Gaussian kernel function; Used to characterize voxels The resonance intensity value.

6. The intelligent vocal training system based on resonance adjustment and breathing feedback according to claim 1, characterized in that: The deep learning model is specifically a multimodal feature fusion model that integrates convolutional neural networks to extract acoustic features and recurrent neural networks to capture temporal dynamics. The deep learning model provides intelligent scoring and suggestions for different vocal registers and breathing patterns, specifically including the following steps: Sc100. Construct the input layer of the deep learning model, expressed as follows: Where: A is used to characterize acoustic feature sequences; B is used to characterize respiratory feature sequences; Sc200. Perform layer normalization and position encoding on the input layer, expressed as follows: in: Used for characterization and features The energy envelope corresponding to the resonance region; Used to characterize and The energy envelope corresponding to the breath region; Sc300. Construct acoustic branches, expressed as follows: in: Used to characterize the process after the first The feature vector at time step t after convolutional processing; Used to characterize the Layer convolution kernel parameters; Used to characterize the number of stacked convolutional layers in a model; Used to characterize the linear rectified activation function; Sc400. Construct respiratory pathways, expressed as follows: Wherein: GRU is used to characterize the gated loop unit operation; Used to characterize the output of the forward GRU; Used to characterize the inverted GRU output; Used to characterize the spliced ​​output of a bidirectional GRU; Sc500. Calculate cross-modal attention, expressed as follows: in: A query matrix used to characterize the resonance channel; A bond matrix used to characterize the respiratory pathway; Used to represent vector dimension; A value matrix used to characterize the respiratory channel; Used to characterize the attention normalization function; Used to characterize the resonance-breathing interaction result calculated by the h-th attention channel; The feedforward network and residuals calculated by Sc600 are expressed by the following formula: in: Used to characterize the output features of the multi-head attention module, serving as a reference baseline for the input of the feedforward module; Used to characterize the output of the feedforward module; Used to characterize at time step The final fusion characteristics; Used to characterize the first-level bias term; Used to characterize the second-layer bias term; ReLU is used to characterize the activation function; LayerNorm is used to characterize layer normalization; Used to characterize the weights of the first fully connected layer; Used to characterize the weights of the second fully connected layer; Sc700 provides intelligent scores and suggestions for different vocal registers and breathing patterns, expressed as follows: , in: The final feature vector used to characterize the fusion of the two channels of resonance and breathing; Used to characterize the weight matrix; Used to characterize the offset of the adjusted score output; Used to characterize the overall score; Used to characterize the attention mapping weight matrix; The attention weight used to characterize the k-th suggestion; Used to characterize the k-th proposal vector; Used to characterize the final proposal vector; Sc800. Calculate the loss function of the deep learning model, specifically expressed as follows: in: Used to characterize the weighting coefficients of the scoring loss; Used to characterize the scoring loss function; Used to characterize the weighting coefficients of the proposed loss; Used to characterize the loss function for the proposed branch; Used to characterize the weight coefficients of the regularization term; Used to characterize the model parameter vector; Used to characterize the total loss function.

7. The intelligent vocal training system based on resonance adjustment and breathing feedback according to claim 1, characterized in that: The manually preset evaluation criteria are used as auxiliary criteria for system output, including formant deviation, tone brightness, breath support, and comprehensive score. The resonance peak deviation is expressed by the following formula: in: Used to characterize the frequency (measured value) of the nth resonant peak; Used to characterize the frequency of the nth resonant peak (reference value); The brightness of the sound quality is expressed by the following formula: in: Used to characterize the frequency of the kth frequency point; Used to characterize the power spectrum value at the k-th frequency point; The breath support level is expressed by the following formula: in: Used to characterize the respiratory sustainment ratio; Used to characterize respiratory pressure stability score; The comprehensive score is expressed as follows: in: Used to characterize resonance weights; Used to characterize sound quality weights; Used to characterize breath weight; Used to characterize the maximum permissible resonance deviation; Used to characterize the maximum frequency.

8. A training method utilizing the intelligent vocal training system based on resonance adjustment and breathing feedback as described in any one of claims 1 to 7, characterized in that: Includes the following steps: S100. Collect basic information about the current training object and initialize the intelligent system; S200. The wearable sensing device (110) is worn by the current training subject; at the same time, the sound acquisition device is set within a preset range of the current user's location; then the wearable sensing device (110) and the sound acquisition device are connected to the training main control system (500); S300. The respiratory data acquisition module (100) reads the respiratory fluctuation signal of the chest and abdomen in real time and automatically adjusts the sampling frequency to adapt to different teaching rhythms. The respiratory fluctuation signal includes the respiratory depth, the respiratory rate, the respiratory rhythm, and the duration; S400. Perform the filtering and normalization processing on the respiratory fluctuation signal collected in step S300; S500. The audio signal of the current training object is collected through the sound acquisition and analysis module (200), and then the acoustic features are extracted through Fourier transform and wave transform; S600. The respiratory and sound feature data obtained after being processed by the respiratory data acquisition module (100) and the sound acquisition and analysis module (200) are uploaded to the training main control system (500) for fusion analysis, and are synchronously transmitted back to the visualization interactive client (600) for display in real time; S700. The training data processed by the training control system (500) is intelligently evaluated in multiple dimensions by the deep learning model built into the intelligent training feedback module (400) and the manually preset evaluation criteria, generating evaluation results for breathing-voice synchronization, resonance concentration, pitch control accuracy, and breath support matching degree; then, waveform diagrams, heat maps, and stability curves are generated and fed back to the visual interactive client (600) and the remote teaching and teaching management platform; S800. Teachers use the remote teaching and teaching management platform to intelligently manage the training process and training data of the current trainees, specifically including real-time remote viewing of the trainees' user data and training data, uploading demonstration content, and issuing tasks for the current trainees; the training master control system (500) automatically generates training progress reports and growth curves to provide trainees with phased feedback and suggestions for the next training step.

9. The training method according to claim 8, characterized in that: The filtering and normalization process described in step S400 specifically includes the following steps: Sd100. Uses a bandpass filter to filter out non-breathing bandwidth; Sd200. Apply Kalman filtering to smooth burst artifacts; Sd300 performs Z-score normalization on the filtered signal to eliminate individual differences and ensure that data from different training subjects are comparable.

10. The training method according to claim 8, characterized in that: The multi-dimensional evaluation results include the degree of synchronization between breathing and vocalization, whether the resonance area is concentrated and stable, the accuracy and stability of pitch control, whether the breath support matches the vocal length and intensity, the breath matching score, and the personalized prompt card for the training status; the output results of the multi-dimensional evaluation results after image processing include a breathing waveform diagram, an acoustic spectrum heat map, a three-dimensional diagram of resonance distribution, and a pitch stability curve.