Intelligent vocal music vocal training device
Through multi-dimensional data analysis and AI algorithms, the problem of insufficient personalization in traditional vocal teaching is solved, accurate evaluation and personalized training of vocal pronunciation are achieved, and the effect of vocal teaching is improved.
Patent Information
- Application Number
- CN202510482274.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional vocal teaching model lacks personalization, making it difficult to identify the pronunciation characteristics and virtual and real voice conversion techniques unique to ancient poems, and the time and space of breathing and vocal data are mismatched, resulting in a reduced teaching effect.
By integrating multi-dimensional data such as oral dynamics and respiratory temperature fields, combined with respiratory muscle group movement models, AI algorithms are used to analyze the matching degree of vocal pronunciation and the melody/tooth rules of ancient poetry and artistic songs, and a personalized training plan is generated.
Accurate evaluation and personalized training of vocal pronunciation are realized, and the correlation between airflow changes and song tone rules is quantified, real-time feedback and action correction are provided, and teaching effect is improved.
Smart Images

Figure CN120388499A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vocal music training, and specifically relates to an intelligent vocal music sound production training device. Background Art
[0002] The traditional vocal music teaching mode refers to the vocal music teaching method based on classical music, which usually includes: basic vocal music training (such as breathing, vocalization, resonance, etc.), song singing (to improve vocal music skills and musical performance ability through singing), music theory (including music score reading, musical rhythm, harmony, tonality, etc.), and stage performance (to improve stage performance ability and musical expressiveness). The traditional vocal music teaching mode pays attention to the training of vocal music skills in its teaching process, emphasizes the purity, elegance and artistry of vocal music, but there are also certain drawbacks: Firstly, there is a lack of personalized teaching. The traditional vocal music teaching mode usually adopts fixed teaching methods and teaching materials, ignoring the individual differences and characteristics of students. The voice texture, technical needs and musical goals of each student are different, and the one-size-fits-all teaching of the traditional mode often fails to meet the needs of different students. Secondly, it attaches importance to skills while ignoring expression.
[0003] In the prior art, such as a vocal music education practice device disclosed in CN116612676A, which includes a training box body, a video recording system, a breathing acquisition part and a vocal cord detection part; the training box body, the front end of which is detachably connected with a display screen, the video recording system, which is arranged on the training box body, the breathing acquisition part, which is worn around the abdomen / chest position of the singer, including a strain sensor that can detect the singer's breathing change and a vibrator for reminding the singer to breathe; the vocal cord detection part, which is worn at the position close to the vocal cords in the singer's throat, including two vibration sensors. This device can help singers improve their vocalization skills, master correct breathing skills and postures, etc., and improve their practice effect through technical means such as the breathing acquisition part, the vocal cord detection part and the video recording system, and has ease of use and easy operation. However, there are also some defects in teaching vocal music to the trainer only by detecting the vibrations in the abdomen and throat parts of the trainer, for example: 1. It is difficult to identify the unique pronunciation characteristics of entering tones and the conversion technique of virtual and real sounds (the air intake control of ±0.3 seconds with the voice interrupted and the breath continuous) in the song (made according to ancient Chinese poems); 2. Only macroscopic vibrations are monitored while the glottis closure degree and nasal cavity resonance energy flow are missing; 3. The breathing and vocalization data are mismatched in time and space, resulting in the failure of the "air - sound - cavity" coordination relationship modeling.
[0004] Therefore, it is necessary to propose an intelligent vocal music sound production training device to solve the above problems. Summary of the Invention
[0005] To solve the above problems, the object of the present invention is to provide an intelligent vocalization training device, which integrates multi-dimensional data such as oral cavity dynamics and respiratory temperature field, combines with the respiratory muscle movement model, analyzes the matching degree between vocal pronunciation and the tune / level and oblique tones rules of ancient poetry art songs, and uses AI algorithms to achieve accurate evaluation and generate personalized training programs.
[0006] To achieve the above object, the technical solution of the present invention is as follows: An intelligent vocalization training device, including a base and an electronically controlled lifting column installed on the top of the base, the top of the electronically controlled lifting column is fixedly connected with a display screen, a first placement rack and a second placement rack are respectively installed on both sides of the display screen, a singing device is placed on the first placement rack, and a belly band is placed on the second placement rack; it also includes a control system, and the display screen, the electronically controlled lifting column, the singing device and the belly band are all signal-connected to the controller system;
[0007] The control system includes:
[0008] An oral cavity monitoring module, used to obtain the opening and closing information of the oral cavity when the user vocalizes, and construct a corresponding three-dimensional oral cavity morphology; it is also used to obtain the change information of the air flow temperature field around the oral cavity when the user vocalizes, and establish an association model between the change information of the air flow temperature field and the tune / level and oblique tones relationship of ancient poetry art songs;
[0009] An abdominal detection module, used to obtain the movement information of the diaphragm and intercostal muscles inside the user's body, and establish a three-dimensional mapping model;
[0010] An AI analysis module, used to perform a joint analysis based on the data of the association model and the three-dimensional mapping model to evaluate the accuracy of the user's singing.
[0011] Furthermore, the oral cavity monitoring module includes an air flow detection unit and a dynamic detection unit;
[0012] The dynamic detection unit uses an infrared laser scanner, which is used to capture the change of the laser reflection path when the oral cavity opens and closes, and capture the movement trajectory of the tongue tip through the diffuse reflection characteristics of the laser to construct a three-dimensional oral cavity morphology; the infrared laser scanner is set on the surface of the singing device;
[0013] The air flow detection unit uses a thermal imager, which is used to capture the change of the air flow temperature around the oral cavity when the user sings a song, and identify and establish a corresponding relationship matrix of the tune / level and oblique tones rules of ancient poetry art songs through the temperature field gradient change mode; the thermal imager is set on the top of the display screen.
[0014] Furthermore, the abdominal detection module includes a biomechanical sensing unit and a data fusion unit;
[0015] The biomechanical sensing unit is used to capture the movement data of the diaphragm and intercostal muscles when the user wears an abdominal belt while singing a song; several piezoresistive flexible sensors and MEMS inertial measurement sub-units are integrated on the inner surface of the abdominal belt;
[0016] The data fusion module is used to dynamically model the movement trajectory of the diaphragm and the cooperative action of the intercostal muscle groups based on the Kalman filter.
[0017] Furthermore, the AI analysis module includes a multi-dimensional comparison unit and a real-time feedback unit;
[0018] The multi-dimensional comparison unit is used to establish a three-dimensional parameter space of acoustic features including pitch, loudness, and timbre;
[0019] The real-time feedback unit is used to synchronously project the three-dimensional oral cavity shape, the airflow temperature field distribution, and the diaphragm movement parameters onto the display screen; the display screen displays in a split-screen mode in real time: the deviation amount between the dynamic waveform of the oral cavity opening and closing angle and the standard song tune template, the spatial superposition mapping of the airflow temperature field pseudo-color cloud map and the rhythm line of the poem or song, and the projection coordinates of the diaphragm displacement trajectory in the three-dimensional acoustic parameter space.
[0020] Furthermore, the real-time feedback unit is also used to integrate a spectral waterfall plot and an error marking layer on the interface of the display screen, and perform accelerated rendering on the spectral waterfall plot and the error marking layer to refresh the morphological data, and trigger the red frame flashing warning and the correction direction vector arrow indication of the target area when a pronunciation error occurs.
[0021] Furthermore, in the airflow detection unit, the spatio-temporal distribution data of the oral cavity opening and closing angle and the airflow temperature field are obtained through a preset sampling frequency, a convolutional neural network classification model of the temperature field gradient change and the level and oblique tones rules of the poem or song is constructed, and then the attention mechanism is used to dynamically weight the contribution degrees of the temperature field features of different pronunciation parts.
[0022] Furthermore, the preset sampling frequency in the airflow detection unit is ≥100 Hz.
[0023] Furthermore, the feature extraction process of the temperature field gradient change is as follows:
[0024] Step 1, establish a motion trajectory model of the vortex center point of the heat flow field;
[0025] Step 2, calculate the temperature gradient covariance matrix of the lip and tooth area and the soft palate area;
[0026] Step 3, analyze the fluctuation spectral characteristics of the temperature field through wavelet transform.
[0027] Furthermore, in the multi-dimensional comparison unit, first establish a comprehensive evaluation index system including pitch accuracy, rhythm, and articulation, and then use federated learning to dynamically adjust personalized parameter thresholds to generate a multi-channel feedback scheme including visual guidance, tactile cues, and auditory correction.
[0028] Furthermore, the feedback scheme includes: a grading prompt mechanism based on the severity of pronunciation errors, constructing a personalized training atlas containing historical error patterns, and 3D interaction for action demonstration and error correction through a virtual teacher avatar.
[0029] Beneficial effects of adopting this solution: 1. Through the synergistic effect of infrared laser scanning and thermal imaging technology, the system converts the abstract relationship between oral cavity dynamics and the rhythm of ancient poetry art songs into a visual model. Laser scanning captures the movement trajectory of the tongue tip, and combined with the analysis of the temperature field gradient, reveals the correlation between air flow changes and the prosody rules of the song, enabling quantitative expression of concepts such as "breath conversion" and "articulation and rhyming" in traditional teaching. The flexible sensor array synchronously monitors the diaphragm displacement and constructs a three-dimensional model of the coordinated action of the respiratory muscles to provide an objective evaluation basis for breath support.
[0030] 2. The split-screen display technology spatially superimposes the dynamic waveform of the oral cavity opening and closing angle, the air flow temperature field distribution, and the diaphragm movement trajectory. Through the visual comparison between the pseudo-color cloud map and the standard template, the pronunciation deviation is intuitively presented. The system adopts dynamic waveform synchronization technology to enable trainees to observe the matching degree of tone trends and prosody rules in real time. When an abnormality in the key pronunciation part is detected, a closed-loop training mechanism of "error perception - immediate feedback - action correction" is established through warning signs in the target area and correction direction guidance. The spatio-temporal correlation analysis of the spectral waterfall diagram and the historical error marking layer helps trainees understand the causal relationship between pronunciation defects and breath control.
[0031] 3. Based on the dynamic parameter adjustment mechanism of federated learning, the system continuously updates the acoustic feature threshold database and dynamically optimizes the feature weights of different pronunciation parts in combination with the attention mechanism. When a continuous deviation of a specific prosody combination is detected, a special training plan including 3D virtual demonstrations is automatically generated, the tongue position adjustment path is decomposed through an anatomical model, and targeted reinforcement is achieved through multi-channel feedback (visual guidance, tactile cues, auditory correction). Brief Description of the Drawings
[0032] Figure 1 It is a flowchart of an embodiment of the intelligent vocalization training device of the present invention;
[0033] Figure 2 It is an axonometric view of an embodiment of the intelligent vocalization training device of the present invention;
[0034] Figure 3 It is an axonometric view of an embodiment of the intelligent vocalization training device of the present invention.
[0035] The reference numerals in the accompanying drawings of the specification include: 1, base; 2, electric control lifting column; 3, display screen; 4, abdominal belt; 5, singing device; 6, thermal imager. Detailed Description of the Invention
[0036] The following is a further detailed description through specific embodiments:
[0037] Embodiment:
[0038] Basically as shown in the accompanying drawings Figure 1 , Figure 2 and Figure 3 : An intelligent vocal music training device, including a base 1 and an electric control lifting column 2 installed on the top of the base 1. The top of the electric control lifting column 2 is fixedly connected with a display screen 3 by screws. During specific use, the height of the electric control lifting column 2 can be adjusted according to the height of the user to ensure that the user's line of sight can better view the content displayed on the display screen 3. At the same time, a first placement rack and a second placement rack are respectively fixedly connected to both sides of the display screen 3 by screws, so that a singing device 5 is placed on the first placement rack and an abdominal belt 4 is placed on the second placement rack, facilitating the taking and putting back of the singing device 5 and the abdominal belt 4 before and after training; it also includes a control system, and the display screen 3, the electric control lifting column 2, the singing device 5 and the abdominal belt 4 are all signal-connected to the controller system.
[0039] When creating art songs based on ancient Chinese poems, usually the melody of the art song is related to the prosody relationship of the ancient Chinese poem itself. In the prior art, by detecting the vibrations of the abdomen and the throat to evaluate the accuracy of the singer, it is difficult to identify the specific articulation characteristics and the conversion techniques of real and false sounds of the song (made according to ancient Chinese poems), and only by detecting the macroscopic vibrations (abdomen and throat), the glottis closure degree and the nasal cavity resonance energy flow will be missing, resulting in a certain deviation in the formulation of subsequent personalized solutions and reducing the teaching and practice effect. For the above problems, the specific solutions of this scheme are as follows:
[0040] The control system includes:
[0041] An oral cavity monitoring module, used to obtain the opening and closing information of the oral cavity when the user vocalizes and construct a corresponding three-dimensional oral cavity shape; it is also used to obtain the change information of the air flow temperature field around the oral cavity when the user vocalizes and establish an association model between the change information of the air flow temperature field and the melody / prosody relationship of the ancient Chinese poem art song.
[0042] An abdominal detection module, used to obtain the movement information of the diaphragm and intercostal muscles inside the user's body and establish a three-dimensional mapping model.
[0043] First, capture the information on the opening and closing of the user's mouth during singing, the data on the change in the airflow temperature field around the mouth, and the movement information of the diaphragm and intercostal muscles inside the user's body. At the same time, when the user sings an art song (based on ancient Chinese poems) through the singing device 5, the singing device 5 will also obtain the voice change information of the singer, that is, collect four-dimensional analysis elements. Thus, through these four-dimensional analysis elements, improve the system's recognition of the unique articulation characteristics of entering tone characters and the technique of converting between solid and virtual sounds (the control of the breathing gap of ±0.3 seconds with the voice interrupted and the breath continuous) in art songs (based on ancient Chinese poems). And by introducing the information on the opening and closing of the user's mouth and the data on the change in the airflow temperature field around the mouth, reduce the possibility of only monitoring macroscopic vibrations and missing the glottal closure degree and the nasal cavity resonance energy flow.
[0044] Specifically, the oral cavity monitoring module includes an airflow detection unit and a dynamic detection unit;
[0045] The dynamic detection unit uses an infrared laser scanner to capture the movement trajectory of the tongue tip by the change in the laser reflection path during the opening and closing of the oral cavity and through the diffuse reflection characteristics of the laser to construct a three-dimensional oral cavity shape; the infrared laser scanner is arranged on the surface of the singing device 5. Among them, the infrared laser scanner can preferably be a high-precision infrared laser scanning array (wavelength 850nm, scanning frequency 120Hz), and 12 groups of VCSEL laser emitters are arranged in a ring on the lip contact surface of the singing device 5. When the user pronounces the finals of the "Jiangyang rhyme", the laser beam forms a diffuse reflection point cloud on the inner wall of the oral cavity, and the change in the reflection path is captured by the built-in CMOS image sensor (resolution 1280×960). For example, when singing the art song "Moon Setting, Crows Crying" based on "Mooring by the Maple Bridge at Night", the system reconstructs the soft palate elevation trajectory with a spatial resolution of 0.1mm. When it is detected that the amount of the tongue tip retraction is less than 1.2mm during the entering tone pronunciation of the character "moon", a warning signal is immediately triggered.
[0046] The airflow detection unit uses a thermal imager 6 to capture the changes in airflow temperature around the user's mouth when singing a song. Through the temperature field gradient change pattern, it identifies and establishes a corresponding relationship matrix of the melody / tone rules of ancient poetry and art songs; the thermal imager 6 is set at the top of the display screen 3. In the airflow detection unit, high-frequency sampling (≥100Hz) is used to obtain the spatiotemporal distribution data of the mouth opening and closing angle and the airflow temperature field, and a convolutional neural network classification model is constructed to classify the temperature field gradient changes and the tone rules of poetry and songs. The attention mechanism is then used to dynamically weight the contribution of the temperature field characteristics of different pronunciation parts. Among them, the thermal imager 6 is preferably an uncooled microbolometer (thermal sensitivity ≤50mK) and can be preferably installed at a 30° depression angle on the top of the display screen 3. For example, when analyzing the line "乱石穿空" (a line in the art song "Nian Nu Jiao·Remembering the Past at Chibi"), the thermal imager 6 captures the temperature field changes in the lip and teeth area at a sampling frequency of 500 Hz, and establishes a temperature gradient vector field model. The system uses a convolutional neural network to identify the temperature characteristics of specific level and oblique combinations. For example, the lip airflow temperature corresponding to the departing tone character "乱" (luan) drops sharply by 3.2°C / s, and a correction vector is generated when the deviation from the standard template exceeds 15%.
[0047] The feature extraction process for temperature field gradient changes is as follows: Step 1, establish the motion trajectory model of the center point of the vortex in the thermal flow field; Step 2, calculate the temperature gradient covariance matrix of the lip and tooth area and the soft palate area; Step 3, analyze the temperature field fluctuation spectrum characteristics through wavelet transform.
[0048] Specifically, the abdomen detection module includes a biomechanical sensing unit and a data fusion unit;
[0049] The biomechanical sensing unit is used to capture diaphragm and intercostal muscle movement data when the user wears the abdominal belt 4 while singing. The inner surface of the abdominal belt 4 is integrated with several piezoresistive flexible sensors and a MEMS inertial measurement subunit. The abdominal belt 4 is preferably made of a multi-layer flexible circuit substrate to ensure that it can adapt to the abdominal contours of most people. The sensitivity of the piezoresistive flexible sensors is preferably 0.5 kPa, while the sampling frequency of the MEMS inertial measurement subunit is preferably 200 Hz. Specifically, when playing long notes in an art song (based on "Yangguan Sandie"), the piezoresistive flexible sensor array monitors the abdominal pressure distribution at a 4×4 mm grid density and combines gyroscope data to construct a diaphragm displacement curve. For example, in the sustained note "Quanjun ganguan guanguan guanguan" ("I urge you to drink another glass of wine"), the system detects that the intercostal muscle synergy coefficient is less than 0.7 at the third second and provides immediate tactile feedback indicating that the breathing fulcrum has moved forward 2 cm.
[0050] The data fusion unit is used to dynamically model the diaphragmatic motion trajectory and the synergistic action of the intercostal muscle groups based on the Kalman filter. An improved strong tracking Kalman filter is adopted to perform spatio-temporal alignment on the multi-source abdominal signals. For example, during the singing of the phrase "As dusk deepens, a clear bugle sounds cold" in the art song (based on "Yangzhou Slow"), the system fuses piezoresistive data (diaphragmatic sinking amount), inertial data (trunk tilt angle), and acoustic features (intensity curve) to establish a respiration-phonation coupling model. When it is detected that the diaphragmatic displacement lags behind the peak sound intensity by 0.3 seconds, the abdominal breathing beat parameters in the training plan are automatically optimized.
[0051] The AI analysis module is used to evaluate the accuracy of the user's singing based on the joint analysis of data from the correlation model and the three-dimensional mapping model, and combines artificial intelligence technology to improve the positive effects of the creation, singing, and teaching of ancient poetry art songs.
[0052] Specifically, the AI analysis module includes a multi-dimensional comparison unit and a real-time feedback unit;
[0053] The multi-dimensional comparison unit is used to establish a three-dimensional parameter space of acoustic features including pitch, intensity, and timbre. At the same time, in the multi-dimensional comparison unit, first, a comprehensive evaluation index system including intonation accuracy, rhythm, and articulation is established, and then the dynamic adjustment of personalized parameter thresholds is realized through federated learning to generate a multi-channel feedback scheme including visual guidance, tactile cues, and auditory correction. The feedback scheme includes: a grading prompt mechanism based on the severity of pronunciation errors, constructing a personalized training map containing historical error patterns, and 3D interaction for action demonstration and error correction through a virtual teacher image.
[0054] The real-time feedback unit is used to synchronously project the three-dimensional oral cavity shape, the airflow temperature field distribution, and the diaphragmatic motion parameters onto the display screen 3; the display screen 3 displays in a split-screen mode in real time: the deviation amount between the dynamic waveform of the oral cavity opening and closing angle and the standard song tune template, the spatial superposition mapping of the airflow temperature field pseudo-color cloud map and the rhythm line of the poetry song, and the projection coordinates of the diaphragmatic displacement trajectory in the three-dimensional acoustic parameter space. Moreover, it is also used to integrate a spectrogram waterfall plot and an error marking layer on the interface of the display screen 3, and perform accelerated rendering on the spectrogram waterfall plot and the error marking layer to refresh the morphological data, and trigger the red frame flashing warning and the correction direction vector arrow indication of the target area when a pronunciation error occurs.
[0055] Taking the art song (produced according to "Reminiscing on the Past at Red Cliff" by Su Shi) as an example, the specific example of the singing evaluation principle of the multi-dimensional comparison unit is as follows:
[0056] When it was detected that the pitch of the word "乱" (falling tone) in "乱石穿空" deviated from the standard template by 12 cents, the federated learning was triggered to dynamically adjust the threshold (the allowable deviation was reduced from ±15 cents to ±8 cents). At the same time, the thermal imager 6 captured the sudden drop in the temperature gradient in the labial and dental area of "岸" (falling tone) at the end of the sentence "惊涛拍岸" (-2.8℃ / ms), which matched the level and oblique tone template with a degree of 92%. When the abdominal belt 4 sensor detected the long tone of "卷起千叠雪", the diaphragm displacement lagged by 0.3 seconds, and the intercostal muscle synergy coefficient was only 0.6 (threshold ≥0.7).
[0057] At this point, in the split-screen mode of the monitor, the left screen displays the waveform of the mouth opening and closing angle (the jaw opening deviation when pronouncing the word "luan" is +3.2mm, and the red waveform exceeds the green template area); the middle screen displays the airflow temperature cloud map (the temperature field and rhythmic line space at the transition point of the level and oblique tones in the sentence "jingtao" are offset by 1.5mm, and the pseudo-color mapping is orange warning); the right screen displays the projection of the diaphragm trajectory in acoustic space (the coordinate point deviates from the threshold boundary of the "breath stability zone", triggering a yellow warning circle). Errors are also marked, and the spectrum waterfall diagram shows that the energy of the final sound of the word "juan" is abnormally attenuated in the 200-400Hz frequency band (the historical error library matches "insufficient abdominal pressure"). The red frame of the target area flashes and a correction arrow pointing to the abdomen is superimposed.
[0058] The feedback system uses a graded prompting system. For a Level 1 error (pitch deviation), the virtual teacher highlights the character "乱" (luan) and plays audio with a standard pitch for comparison. For a Level 2 error (respiratory lag), the abdominal belt 4 applies a 5Hz vibration prompt to guide the timing of diaphragmatic depression. Combined with historical data, a specialized exercise for "abdominal pressure support for falling tone characters" was generated, consisting of three sets of singing along with a dense falling tone section of the art song. Finally, the virtual teacher breaks down and demonstrates the tongue root retraction movement in the phrase "惊涛拍岸" (the three-dimensional tongue position model gradually changes from a red incorrect position to a green standard position).
[0059] The above is only an embodiment of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme is not described in detail here. Ordinary technicians in the field are aware of all common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the field can improve and implement this scheme in combination with their own abilities under the inspiration given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. An intelligent vocalization training device, comprising a base (1) and an electrically controlled lifting column (2) installed on the top of the base (1), characterized in that, A display screen (3) is fixedly connected to the top of the electric control lifting column (2). A first placement rack and a second placement rack are respectively installed on both sides of the display screen (3). A singing device (5) is placed on the first placement rack, and a belly band (4) is placed on the second placement rack. It also includes a control system. The display screen (3), the electric control lifting column (2), the singing device (5) and the belly band (4) are all signal-connected to the controller system; The control system includes: An oral cavity monitoring module, which is used to obtain the opening and closing information of the oral cavity when the user is singing, and construct a corresponding three-dimensional oral cavity shape. It is also used to obtain the change information of the airflow temperature field around the oral cavity when the user is singing, and establish an association model between the change information of the airflow temperature field and the tune / level and oblique tones relationship of ancient poetry art songs; An abdominal detection module, which is used to obtain the movement information of the diaphragm and intercostal muscles inside the user's body, and establish a three-dimensional mapping model; An AI analysis module, which is used to conduct a joint analysis based on the data of the association model and the three-dimensional mapping model to evaluate the accuracy of the user's singing.
2. The intelligent vocalization training device according to claim 1, wherein: The oral cavity monitoring module includes an airflow detection unit and a dynamic detection unit; The dynamic detection unit uses an infrared laser scanner, which is used to capture the movement trajectory of the tongue tip through the change of the laser reflection path when the oral cavity opens and closes, and capture the movement trajectory of the tongue tip through the diffuse reflection characteristics of the laser to construct a three-dimensional oral cavity shape. The infrared laser scanner is arranged on the surface of the singing device (5); The airflow detection unit uses an infrared thermal imager (6), which is used to capture the change of the airflow temperature around the oral cavity when the user is singing a song, and identify and establish a corresponding relationship matrix between the tune / level and oblique tones rules of ancient poetry art songs through the temperature field gradient change mode. The infrared thermal imager (6) is arranged on the top of the display screen (3).
3. The intelligent vocalization training device according to claim 1, wherein: The abdominal detection module includes a biomechanical sensing unit and a data fusion unit; The biomechanical sensing unit is used to capture the movement data of the diaphragm and intercostal muscles when the user is singing a song after wearing the belly band (4). A number of piezoresistive flexible sensors and MEMS inertial measurement sub-units are integrated on the inner surface of the belly band (4); The data fusion unit is used to dynamically model the movement trajectory of the diaphragm and the synergistic effect of the intercostal muscle groups based on the Kalman filter.
4. The intelligent vocalization training device according to claim 1, wherein: The AI analysis module includes a multi-dimensional comparison unit and a real-time feedback unit; The multi-dimensional comparison unit is used to establish a three-dimensional parameter space of acoustic features including pitch, loudness and timbre; The real-time feedback unit is used to synchronously project the three-dimensional oral cavity shape, the airflow temperature field distribution and the diaphragm movement parameters onto the display screen (3). The display screen (3) displays in a split-screen mode in real time: the deviation amount between the dynamic waveform of the oral cavity opening and closing angle and the standard song tune template, the spatial superposition mapping of the airflow temperature field pseudo-color cloud map and the poem song rhythm line, and the projection coordinates of the diaphragm displacement trajectory in the three-dimensional space of acoustic parameters.
5. The intelligent vocalization training device according to claim 4, wherein: The real-time feedback unit is also used to integrate a spectral waterfall diagram and an error marking layer on the interface of the display screen (3), and perform accelerated rendering on the spectral waterfall diagram and the error marking layer to refresh the morphological data, and trigger the red frame flashing warning and correction direction vector arrow indication of the target area when there is a pronunciation error.
6. The intelligent vocalization training device according to claim 5, wherein: In the airflow detection unit, spatio-temporal distribution data of the oral opening and closing angle and the airflow temperature field are obtained through a preset sampling frequency, a convolutional neural network classification model of the temperature field gradient change and the tonal rules of poetry and songs is constructed, and then an attention mechanism is used to dynamically weight the contribution degrees of the temperature field features of different pronunciation parts.
7. The intelligent vocalization training device according to claim 6, characterized in that: The preset sampling frequency in the airflow detection unit is ≥ 100 Hz.
8. The intelligent vocalization training device according to claim 6, wherein: The feature extraction process of the temperature field gradient change is as follows: Step 1, establish a motion trajectory model of the vortex center point of the heat flow field; Step 2, calculate the temperature gradient covariance matrix of the lip and tooth area and the soft palate area; Step 3, analyze the temperature field fluctuation spectrum characteristics through wavelet transform.
9. The intelligent vocalization training device according to claim 8, characterized in that: In the multi-dimensional comparison unit, first establish a comprehensive evaluation index system including pitch accuracy, rhythm, and articulation, and then realize the dynamic adjustment of personalized parameter thresholds through federated learning to generate a multi-channel feedback scheme including visual guidance, tactile prompts, and auditory correction.
10. The intelligent vocalization training device according to claim 9, characterized in that: The feedback scheme includes: a grading prompt mechanism based on the severity of pronunciation errors, constructing a personalized training atlas containing historical error patterns, and 3D interaction for action demonstration and error correction through a virtual teacher image.
Citation Information
Patent Citations
Vocal music education practice device
CN116612676A