English pronunciation practice auxiliary device
Through the English pronunciation exercise auxiliary equipment integrating the radio and camera components, combining voice and visual analysis, multimodal feedback and cartoon interaction are provided, the existing equipment cannot accurately analyze and lack of fun, and the learning effect is improved.
Patent Information
- Application Number
- CN202510328587.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing English pronunciation practice equipment cannot accurately analyze the dynamic characteristics of pronunciation and facial features at the same time, which lacks fun and interactivity, resulting in low learning efficiency.
Design an English pronunciation exercise auxiliary device, integrates radio and camera components, combines speech spectrum analysis and visual algorithms, and provides multimodal feedback and entertainment interaction through vibration feedback and cartoonized situational interaction, improving learning interest and accuracy.
It realizes accurate analysis and interesting interaction of pronunciation, improves children's concentration and learning effect in English learning, and stimulates practitioners' initiative and interest in exploration.
Smart Images

Figure CN120340322A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of language learning, and particularly to an auxiliary device for English pronunciation practice. Background Art
[0002] With the development of globalization and the increase in international exchanges, English, as an international common language, has become increasingly important in the fields of education and occupation. However, for learners whose native language is not English, especially children, English pronunciation learning is often a difficult point. Inaccurate pronunciation not only affects the clarity of language expression but also may hinder the improvement of language communication ability. Therefore, how to improve the accuracy of learners' English pronunciation through effective teaching means has become an urgent problem to be solved in the field of education.
[0003] Currently, there are some English pronunciation practice devices and methods on the market, such as language learning software or intelligent hardware devices that utilize speech recognition technology. These devices mainly provide pronunciation feedback or scoring functions through the recording and analysis of audio signals. However, such devices usually only rely on the analysis of speech data and ignore the importance of facial dynamics (such as mouth opening and closing, tongue position, etc.) during the pronunciation process, and these dynamic features directly affect the quality of pronunciation. In addition, the interaction methods of most devices are single and lack interest. Especially for child users, the learning process is prone to become boring, thus reducing the learning effect.
[0004] In view of the above problems, the existing technology urgently needs a pronunciation practice device that can combine speech and visual dynamic feature analysis, and at the same time has high interest and interactivity, so as to improve the learning interest and participation of learners and provide more comprehensive pronunciation training guidance.
[0005] To solve the above technical problems, the present invention provides an auxiliary device for English pronunciation practice. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention provides an auxiliary device for English pronunciation practice, which solves the problems that the existing English pronunciation practice devices cannot accurately analyze speech and facial dynamic features simultaneously, and lack of interest and interactivity, resulting in low learning efficiency.
[0007] To achieve the above object, the present invention is realized through the following technical solutions: An auxiliary device for English pronunciation practice, including a storage board, on the surface of which a sound collection component and a camera component are provided. An entertainment area is opened on the upper surface of the storage board, and an entertainment mechanism is arranged in the entertainment area. The entertainment mechanism includes a vibration component and a moving component;
[0008] The storage board is connected to a display board through a rotating shaft, and a display screen is installed on the outer wall of the display board. The display screen is electrically connected to a speech-assisted learning system.
[0009] Preferably, the sound collection component includes a first storage groove which is opened on the upper surface of the storage board. An adjusting seat is rotatably arranged on the inner wall of the first storage groove. One end of a carbon steel spring is installed at the end of the adjusting seat, and the other end of the carbon steel spring is installed with a sound collection tube.
[0010] Preferably, the camera component includes a second storage groove which is opened on the upper surface of the storage board. One end of a folding frame is rotatably arranged on the inner wall of the second storage groove, and the other end of the folding frame is rotatably installed with a camera.
[0011] Preferably, the vibration component includes multiple groups of embedding grooves, and each group of embedding grooves is opened at the bottom of the entertainment area. Multiple groups of guiding grooves are opened inside the storage board, and each guiding groove communicates with the embedding groove. A limiting convex block is installed at the connection between the embedding groove and the guiding groove. A magnetic ring is arranged in the guiding groove, and a push rod is arranged inside the magnetic ring. An attracting piece is installed at the top end of the push rod, and a clamping ring is installed at the bottom end of the push rod. The outer diameter of the clamping ring is larger than the inner diameter of the limiting convex block.
[0012] Preferably, the moving component includes multiple installation grooves, and each installation groove is opened on the surface of the entertainment area. A motor is installed at one end of the inner wall of the installation groove, and a reciprocating lead screw is installed at the driving end of the motor. A moving seat is threaded on the outer wall of the reciprocating lead screw, and a moving piece is installed at the top of the moving seat.
[0013] Preferably, multiple groups of the embedding grooves and multiple of the installation grooves are arranged alternately.
[0014] Preferably, a third storage groove is opened at the top of the storage board, and an operating pen is arranged in the third storage groove.
[0015] Preferably, the voice-assisted learning system includes:
[0016] A data acquisition module, which is used to collect the voice data and visual data of the practitioner in real time as the basis for pronunciation analysis;
[0017] A data analysis module, which is used to compare the collected data with the standard pronunciation model through voice spectrum analysis and visual algorithms to identify the types of pronunciation problems;
[0018] A vibration feedback module, which is used to provide tactile feedback according to the system analysis results and prompt the practitioner whether the pronunciation is correct through vibrations of different frequencies;
[0019] A voice feedback module, which is used to provide voice prompts according to the pronunciation problems to guide the practitioner to adjust the pronunciation posture;
[0020] A visualization feedback module, which is used to dynamically display the lip comparison diagram, adjustment animation and reward animation on the display screen to visually guide the practitioner to adjust the pronunciation;
[0021] An entertainment interaction module, which is used to simulate a cartoonized scenario during the practice process and attract the attention of the practitioner through the dynamic actions of the attracting part and the moving part;
[0022] A learning record and summary module, which is used to record the pronunciation data and problem types of each practice, generate a learning report, and display the user's learning progress and trends.
[0023] Preferably, the data analysis module includes:
[0024] An audio processing unit, which is used to collect the pronunciation data of the practitioner, perform sampling, preprocessing, feature extraction and spectrum analysis, extract pitch, tone and duration features, and perform matching calculations with the standard pronunciation model;
[0025] A visual processing unit, which is used to capture the dynamic movements of the practitioner's mouth shape, tongue position and Adam's apple, extract the key point coordinates and dynamic features, calculate the mouth opening and closing ratio and the tongue position, and judge the difference from the standard model;
[0026] A data fusion and model comparison unit, which is used to perform weighted processing on the audio and visual features, perform matching calculations with the standard pronunciation model, identify the pronunciation problems of the practitioner and locate the deviation content.
[0027] Preferably, the specific calculation of the mouth opening and closing ratio and the tongue position in the visual processing unit includes calculating the mouth opening and closing ratio, the height of the tongue position and the relative position of the mouth shape and the tongue. The formulas are as follows:
[0028] The formula for calculating the mouth opening and closing ratio, which is used to evaluate the degree of mouth opening and closing, is calculated by the distance between the key points of the upper and lower lips and the key points of the corners of the mouth. Specifically:
[0029]
[0030] Among them, R is the mouth opening and closing ratio, p uppe_lip is the coordinate of the midpoint of the upper lip, p lower_lip is the coordinate of the midpoint of the lower lip, p left_corner is the coordinate of the left corner of the mouth, p right_corner is the coordinate of the right corner of the mouth, ||·|| is the Euclidean distance, indicating the straight-line distance between two points;
[0031] The formula for calculating the height of the tongue position, which is used to evaluate the vertical position of the tongue in the mouth, is calculated by the distance between the key points of the tongue and the key points at the bottom of the mouth. Specifically:
[0032] H = ∥p tongu_center -p mout_bottom ∥
[0033] Among them, H is the height of the tongue position, p tongu_center is the coordinate of the center point of the tongue, p mout_bottomare the coordinates of the center point at the bottom of the mouth;
[0034] The calculation formula for the relative position of the mouth shape and the tongue is used to further evaluate the relative dynamics of the tongue position within the mouth. Specifically:
[0035]
[0036] Among them, Relative_Position is the relative position of the tongue and the opening / closing of the mouth shape, H is the height of the tongue position, and R is the mouth opening / closing ratio.
[0037] The present invention provides an auxiliary device for English pronunciation practice, having the following beneficial effects:
[0038] 1. Through the collaborative work of the audio processing unit and the visual processing unit in the data acquisition module of the present invention, the pronunciation data and facial dynamic features of the practitioner can be collected in real time. The audio processing unit performs sampling, preprocessing, spectral analysis, and dynamic time warping algorithms to accurately extract the pitch, tone, and duration features of the practitioner, and performs matching calculations with the standard pronunciation model. At the same time, the visual processing unit realizes the dynamic monitoring of the pronunciation posture by calculating the mouth opening / closing ratio, the height of the tongue position, and the relative position of the mouth shape and the tongue. Based on the design of a formula-based scientific algorithm, the system can accurately identify the specific types of pronunciation problems, such as insufficient mouth opening, low tongue height, etc., thus providing an accurate basis for subsequent corrective feedback.
[0039] 2. The present invention effectively improves the interest and participation of the practitioner through the entertainment interaction module. In the entertainment area of the device, using the linkage mechanism of the vibration component and the moving component, a cartoon-like scenario is simulated, such as the dynamic interaction of "a little rabbit stealing carrots". The attracting member is vibrated by the magnetic ring drive, and the grabbing action is completed by the motor and the moving member, which is linked in real time with the pronunciation state of the practitioner. The system controls the behavior of the entertainment mechanism according to the pronunciation situation of the practitioner, making the practice process full of fun and interactivity, which is especially suitable for the language learning needs of children. This immersive experience not only improves the learning concentration but also stimulates the initiative and exploration interest of the practitioner. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a perspective view of the device of the present invention in the use state;
[0041] Figure 2 is a schematic diagram of the device of the present invention unfolded;
[0042] Figure 3 is a schematic diagram of the internal structure of the embedding groove of the present invention;
[0043] Figure 4 is Figure 3 the enlarged view at A in
[0044] Figure 5 Schematic diagram of the internal structure of the installation groove of the present invention;
[0045] Figure 6 is Figure 5 Enlarged view at position B in
[0046] Figure 7 Framework diagram of the voice-assisted learning system of the present invention;
[0047] Figure 8 Framework diagram of the data analysis module of the present invention.
[0048] Among them, 1. Storage board; 2. Rotating shaft; 3. Display board; 4. Display screen; 5. First storage groove; 6. Adjusting seat; 7. Carbon steel spring; 8. Microphone; 9. Second storage groove; 10. Folding bracket; 11. Camera; 12. Entertainment area; 13. Embedded groove; 14. Guide groove; 15. Magnetic ring; 16. Thumb rod; 17. Snap ring; 18. Limit bump; 19. Attracting member; 20. Installation groove; 21. Motor; 22. Reciprocating lead screw; 23. Moving seat; 24. Moving member; 25. Third storage groove; 26. Operating pen. Specific embodiments
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0050] Please refer to the attached Figure 1 - attached Figure 2 , the embodiment of the present invention provides an English pronunciation practice assistance device, including a storage board 1, a sound collection component and a camera component are arranged on the surface of the storage board 1, an entertainment area 12 is opened on the upper surface of the storage board 1, and an entertainment mechanism is arranged in the entertainment area 12, which is used to simulate a cartoonized scenario and attract children's attention through dynamic actions, increasing the fun of the practice process. The storage board 1 is connected to a display board 3 through a rotating shaft 2, a display screen 4 is installed on the outer wall of the display board 3, and the display screen 4 is electrically connected to a voice-assisted learning system, which is used to collect and analyze the pronunciation data of the practitioner and provide real-time feedback to control the entertainment mechanism and guide the pronunciation training by voice.
[0051] Specifically, the device of the present invention adopts a notebook folding design. A foldable structure is formed through the rotational connection between the storage board and the display board, which is convenient for carrying and protecting internal components. The surface of the storage board is integrated with a sound collection component and a camera component, which are used to collect the voice data and facial dynamics of the practitioner. The sound collection component uses an adjustable structure to collect the pronunciation characteristics of the practitioner, while the camera component provides morphological data for pronunciation analysis by capturing the dynamics of the practitioner's lip shape, tongue position, and laryngeal prominence in real time.
[0052] Entertainment institutions are arranged in the entertainment area to attract children's attention through cartoonized dynamic simulation. When the practitioner pronounces, the voice-assisted learning system controls the actions of the entertainment institutions by analyzing the voice data to simulate cartoonized scenarios, such as dynamic grasping or lifting actions, to stimulate children's interest in participation. The display screen shows the pronunciation status and improvement suggestions of the practitioner in real time, and adjusts the pronunciation through voice guidance to form a learning loop.
[0053] Please refer to the appendix Figure 1 - appendix Figure 2 , the sound collection component includes a first storage groove 5, the first storage groove 5 is opened on the upper surface of the storage board 1, a regulating seat 6 is rotatably arranged on the inner wall of the first storage groove 5, one end of a carbon steel spring 7 is installed at the end of the regulating seat 6, and the other end of the carbon steel spring 7 is installed with a sound collecting cylinder 8.
[0054] Specifically, the first storage groove provides structural protection and adjustable functions. The regulating seat is installed on the inner wall of the first storage groove and can rotate freely to adjust the angle and position when the practitioner uses it. One end of the carbon steel spring is connected to the regulating seat, and the other end is connected to the sound collecting cylinder. Using the elastic support and stabilizing effect, it ensures that the sound collecting cylinder remains fixed at different adjustment angles. When the practitioner is in pronunciation training, according to the position of their own mouth, the orientation of the sound collecting cylinder can be changed through the regulating seat to align it with the mouth area to ensure the accuracy of voice signal collection. After use, the regulating seat and the carbon steel spring retract the sound collecting cylinder into the storage groove to ensure the cleanliness and safety of the device. The sound collecting cylinder is used to collect the voice signal of the practitioner and transmit it to the voice-assisted learning system for analysis.
[0055] Please refer to the appendix Figure 1 - appendix Figure 2 , the camera component includes a second storage groove 9, the second storage groove 9 is opened on the upper surface of the storage board 1, one end of a folding frame 10 is rotatably arranged on the inner wall of the second storage groove 9, and the other end of the folding frame 10 is rotatably installed with a camera 11.
[0056] Specifically, the camera assembly is arranged on the surface of the storage board. The second storage groove is used to accommodate and protect the structure of the camera and the folding rack. One end of the folding rack is fixed to the inner wall of the second storage groove and can rotate around the axis point, and the other end is connected to the camera. When in use, the practitioner rotates out the folding rack and adjusts the angle of the camera so that it faces the mouth and laryngeal prominence areas. The camera collects data on the opening and closing of the practitioner's mouth shape, tongue movement, and laryngeal dynamics, which serves as an important basis for pronunciation analysis. After use, the folding rack and the camera are rotated back into the second storage groove to ensure the cleanliness and safety of the device.
[0057] Please refer to the appendix Figure 3 - appendix Figure 6 , the entertainment mechanism includes a vibration component and a moving component. The vibration component is used to achieve the vibration of the ejector rod 16 through electromagnetic drive, generating tactile feedback of different frequencies and intensities. The moving component is used to drive the moving part to simulate a grasping action and cooperate with the vibration component to complete dynamic interaction;
[0058] The vibration component includes multiple groups of embedding grooves 13. Each group of embedding grooves 13 is opened at the bottom of the entertainment area 12. Multiple groups of guiding grooves 14 are opened inside the storage board 1. Each guiding groove 14 communicates with the embedding groove 13. A limiting convex block 18 is installed at the connection of the embedding groove 13 and the guiding groove 14. A magnetic ring 15 is arranged in the guiding groove 14. An ejector rod 16 is arranged in the magnetic ring 15. An attracting part 19 is installed at the top of the ejector rod 16. A snap ring 17 is installed at the bottom of the ejector rod 16. The outer diameter of the snap ring 17 is larger than the inner diameter of the limiting convex block 18.
[0059] The moving component includes multiple installation grooves 20. Each installation groove 20 is opened on the surface of the entertainment area 12. One end of the inner wall of the installation groove 20 is installed with a motor 21. The driving end of the motor 21 is installed with a reciprocating lead screw 22. A moving seat 23 is threaded on the outer wall of the reciprocating lead screw 22. A moving part 24 is installed on the top of the moving seat 23.
[0060] Multiple groups of embedding grooves 13 and multiple installation grooves 20 are arranged alternately.
[0061] Specifically, the entertainment mechanism is composed of a vibration component and a moving component. By coordinating actions, it simulates a cartoon-like dynamic scene to achieve the entertainment interaction function. The vibration component uses the magnetic ring to drive the ejector rod to generate vibration, causing the attracting part to rise from the embedding groove to simulate the dynamics of an object. The moving component drives the reciprocating lead screw to rotate through the motor, pushing the moving part to move along a fixed track and cooperating with the attracting part of the vibration component to complete the simulation of dynamic grasping.
[0062] In the vibration assembly, the buried groove communicates with the guiding groove, and the magnetic ring is installed in the guiding groove. After being powered on, the magnetic field generated by the magnetic ring drives the ejector rod to vibrate, and the attracting member rises accordingly. The snap ring at the bottom of the ejector rod cooperates with the limiting bump to limit the vibration range of the ejector rod, ensuring the movement stability and safety of the attracting member. The moving assembly drives the reciprocating lead screw to rotate through a motor, and the threaded structure on the lead screw pushes the moving seat to reciprocate along the track. The moving member on the top of the moving seat imitates the grasping action and forms a dynamic interaction with the attracting member of the vibration assembly.
[0063] Please refer to the appendix Figure 1 - appendix Figure 2 , a storage groove III 25 is provided at the top of the storage board 1, and an operating pen 26 is arranged in the storage groove III 25.
[0064] Specifically, the storage groove III is used to store the operating pen, providing a convenient interaction tool for the device, and ensuring that the operating pen is properly stored when not in use. The operating pen is placed inside the storage groove III, and through the precise dimensions and fixed design of the groove body, it is ensured that it will not fall off or be damaged during the transportation or storage of the device. When in use, the practitioner takes out the operating pen from the storage groove III and interacts with the device through it to complete tasks such as mode selection and function operation.
[0065] Please refer to the appendix Figure 7 , the voice-assisted learning system includes:
[0066] A data acquisition module for real-time acquisition of the voice data and visual data of the practitioner as the basis for pronunciation analysis;
[0067] A data analysis module for comparing the acquired data with the standard pronunciation model through voice spectrum analysis and visual algorithms to identify the types of pronunciation problems;
[0068] A vibration feedback module for providing tactile feedback according to the system analysis results and prompting the practitioner whether the pronunciation is correct through vibrations of different frequencies;
[0069] A voice feedback module for providing voice prompts according to the pronunciation problems to guide the practitioner to adjust the pronunciation posture;
[0070] A visualization feedback module for dynamically displaying the mouth shape comparison diagram, adjustment animation, and reward animation through the display screen 4 to intuitively guide the practitioner to adjust the pronunciation;
[0071] An entertainment interaction module for simulating a cartoonized scenario during the practice process and attracting the practitioner's attention through the dynamic actions of the attracting member 19 and the moving member 24;
[0072] A learning record and summary module for recording the pronunciation data and problem types of each practice, generating a learning report, and showing the user's learning progress and trend.
[0073] Specifically, in this embodiment, the data acquisition module includes an audio acquisition and a visual acquisition unit, which are used to collect the voice and visual data of the practitioner in real time.
[0074] Audio acquisition:
[0075] The voice signal of the practitioner is collected by the microphone 8 and digitized with a sampling frequency of 16 kHz and a quantization depth of 16 bits.
[0076] Noise preprocessing:
[0077] A band-pass filter is used to eliminate background noise, and the passband range of the filter is 300 Hz - 3400 Hz.
[0078] Apply mean filtering to smooth the signal.
[0079] Frame segmentation and window function:
[0080] The voice signal x[n] is divided into frames, each frame is 20 ms long, and the frame shift is 10 ms.
[0081] Apply the Hamming window w[n] to reduce spectral leakage: x w [n] = x[n] · w[n]
[0082] Short-time Fourier transform (STFT):
[0083] Perform STFT on each frame of the signal to obtain spectral features:
[0084]
[0085] Among them, X(k,m) represents the spectrum of the m-th frame, and N is the frame length.
[0086] Visual acquisition:
[0087] Use the camera 11 to collect the video stream in real time and capture the dynamics of the lip shape, tongue position, and Adam's apple.
[0088] Preprocessing:
[0089] Gray-scale the image and perform noise suppression.
[0090] Crop the face region through the Haar classifier or the MTCNN model.
[0091] Key point detection:
[0092] Use the MediaPipeFaceMesh algorithm to extract the key points of the mouth (such as the corners of the mouth, the midpoint of the upper lip, and the midpoint of the lower lip):
[0093] P = {p1, p2, …, p k}
[0094] Where each point pk =(x k , y k ) represents a two-dimensional coordinate.
[0095] In this embodiment, the data analysis module compares the collected data with the standard model through voice spectrum analysis and visual algorithms to identify the types of pronunciation problems.
[0096] Audio analysis:
[0097] Extract the following audio features:
[0098] Fundamental frequency (F0): Use the autocorrelation function method to extract the main frequency of the speech.
[0099] Formants: Estimate the formant frequency F k :
[0100] F k = argmax f |X(f)|
[0101] Dynamic Time Warping (DTW) matching:
[0102] Perform time alignment on the audio feature sequence:
[0103] D(i, j) = d(i, j) + min{D(i - 1, j), D(i, j - 1), D(i - 1, j - 1)}
[0104] Matching score calculation:
[0105]
[0106] Visual analysis:
[0107] Mouth opening degree:
[0108] Calculate the mouth opening ratio:
[0109]
[0110] where R is the mouth opening ratio, p uppe_lip is the coordinate of the midpoint of the upper lip, p lower_lip is the coordinate of the midpoint of the lower lip, p left_corner is the coordinate of the left corner of the mouth, p right_corner is the coordinate of the right corner of the mouth, ||·|| is the Euclidean distance, representing the straight-line distance between two points.
[0111] Tongue position height:
[0112] Calculate the distance between the center point of the tongue and the bottom of the mouth:
[0113] H = ∥ptongu_center -p mout_bottom ∥
[0114] Among them, H is the height of the tongue position, and p tongu_center is the coordinate of the center point of the tongue, and p mout_bottom is the coordinate of the center point of the bottom of the mouth.
[0115] Data fusion:
[0116] Perform weighted fusion on audio and visual features:
[0117] F combined = w1·F audio + w2·F visual
[0118] Among them, w1 and w2 are weight parameters.
[0119] In this embodiment, the vibration feedback module drives the ejector rod 16 through the magnetic coil 15 to generate vibration prompts of different frequencies to indicate the types of pronunciation problems.
[0120] Vibration frequency classification:
[0121] Low-frequency vibration (30 - 50 Hz): Indicates correct pronunciation.
[0122] Medium-frequency vibration (50 - 70 Hz): Indicates a slight deviation.
[0123] High-frequency vibration (70 - 100 Hz): Indicates a significant deviation.
[0124] Drive control:
[0125] Use a PWM signal to adjust the magnetic coil current and control the vibration frequency and amplitude of the ejector rod:
[0126] I = I max ·sin(2πft)
[0127] Among them, f is the vibration frequency.
[0128] In this embodiment, the voice feedback module plays voice prompts through a speaker to guide the practitioner to adjust their pronunciation.
[0129] Voice prompt content:
[0130] Regarding tongue position problems: "Lift your tongue and touch your upper teeth."
[0131] Regarding mouth shape problems: "Open your mouth wider."
[0132] Playback logic:
[0133] The prompt content is dynamically generated according to the data analysis results and is synchronized with the vibration feedback.
[0134] In this embodiment, the visualization feedback module displays the real-time mouth shape comparison diagram and the dynamic adjustment animation through the display screen 4.
[0135] Mouth shape comparison:
[0136] The standard mouth shape model is displayed on the left, and the real-time mouth shape is displayed on the right.
[0137] The deviation area is marked in red, and the correct area is marked in green.
[0138] Dynamic adjustment animation:
[0139] Displays the movement path of the tongue and the dynamic opening and closing of the mouth shape.
[0140] In this embodiment, the entertainment interaction module simulates a cartoonized scenario through the attracting member 19 and the moving member 24 to attract the attention of the practitioner.
[0141] Trigger mechanism:
[0142] When pronunciation is detected, the magnetic ring drives the ejector rod to raise the attracting member, and the motor 21 drives the moving member to simulate a grasping action.
[0143] Scene design:
[0144] Cooperate with voice prompts, such as "The little rabbit is going to steal the radish. Hurry up and stop it with the correct pronunciation!"
[0145] In this embodiment, the learning record and summary module records the pronunciation data and problem types of each practice and generates a learning report.
[0146] Record content:
[0147] Pronunciation accuracy rate, problem type, practice duration.
[0148] Report generation:
[0149] Automatically generate a user learning trend chart, including the change in pronunciation accuracy rate and the classification of major problems.
[0150] Please refer to the appendix Figure 8 , the data analysis module includes:
[0151] An audio processing unit for collecting the pronunciation data of the practitioner, performing sampling, preprocessing, feature extraction, and spectrum analysis, extracting pitch, tone, and duration features, and performing matching calculations with the standard pronunciation model;
[0152] A visual processing unit for capturing the mouth shape, tongue position, and laryngeal knot dynamics of the practitioner, extracting the key point coordinates and dynamic features, calculating the mouth opening and closing ratio and the tongue position, and judging the difference from the standard model;
[0153] The data fusion and model comparison unit is used to perform weighted processing on audio and visual features, perform matching calculations with the standard pronunciation model, identify the pronunciation problems of the practitioner and locate the deviation content.
[0154] Specifically, in the visual processing unit, calculating the mouth opening and closing ratio and the tongue position includes calculating the mouth opening and closing ratio, the height of the tongue position, and the relative position between the mouth shape and the tongue. The formulas are as follows:
[0155] The formula for calculating the mouth opening and closing ratio is used to evaluate the degree of mouth opening and closing. It is calculated by the distance between the key points of the upper and lower lips and the key points of the corners of the mouth. Specifically:
[0156]
[0157] where R is the mouth opening and closing ratio, p uppe_lip is the coordinate of the midpoint of the upper lip, p lower_lip is the coordinate of the midpoint of the lower lip, p left_corner is the coordinate of the left corner of the mouth, p right_corner is the coordinate of the right corner of the mouth, and ||·|| is the Euclidean distance, representing the straight-line distance between two points;
[0158] The formula for calculating the height of the tongue position is used to evaluate the vertical position of the tongue in the mouth. It is calculated by the distance between the key point of the tongue and the key point of the bottom of the mouth. Specifically:
[0159] H = ∥p tongu_center - p mout_bottom ∥
[0160] where H is the height of the tongue position, p tongu_center is the coordinate of the center point of the tongue, p mout_bottom is the coordinate of the center point of the bottom of the mouth;
[0161] The formula for calculating the relative position between the mouth shape and the tongue is used to further evaluate the relative dynamics of the tongue position in the mouth. Specifically:
[0162]
[0163] where Relative_Position is the relative position between the tongue and the mouth opening and closing, H is the height of the tongue position, and R is the mouth opening and closing ratio.
[0164] Specifically, in this embodiment, the audio processing unit is used to collect the pronunciation data of the practitioner, complete sampling, preprocessing, feature extraction, and spectrum analysis, and perform matching calculations with the standard pronunciation model.
[0165] Speech signal preprocessing:
[0166] The pronunciation data is discretized at a sampling rate of 16 kHz to generate a digital signal sequence:
[0167]
[0168] Use a band - pass filter to remove ambient noise and retain the frequency components from 300 Hz to 3400 Hz.
[0169] Feature extraction:
[0170] Short - time Fourier transform (STFT):
[0171]
[0172] X(k,m): The spectrum of the m - th frame.
[0173] w[n]: Window function, R: Frame shift length.
[0174] Formant extraction:
[0175] F k = argmax f |X(f)|
[0176] F k : The k - th formant frequency.
[0177] Matching calculation:
[0178] Use the dynamic time warping (DTW) algorithm to compare the extracted features with the standard pronunciation model:
[0179] D(i,j)= d(i,j)+ min{D(i - 1,j),D(i,j - 1),D(i - 1,j - 1)}
[0180] Matching score calculation formula:
[0181]
[0182] In this embodiment, the visual processing unit captures the dynamic mouth shape, tongue position, and laryngeal prominence of the practitioner through a camera, calculates the mouth opening ratio, tongue position height, and the relative position between the mouth shape and the tongue, and judges the difference between the practitioner and the standard model.
[0183] Mouth opening ratio calculation formula:
[0184] The mouth opening ratio is used to evaluate the degree of mouth opening. The calculation formula is:
[0185]
[0186] Among them, R is the mouth opening ratio, p uppe_lip is the coordinate of the mid - point of the upper lip, p lower_lip is the coordinate of the mid - point of the lower lip, p left_corner is the coordinate of the left corner of the mouth, pright_corner is the coordinate of the right corner of the mouth;
[0187] ||·|| is the Euclidean distance, representing the straight-line distance between two points:
[0188] Calculate the straight-line distance between two points:
[0189]
[0190] Tongue position height calculation formula:
[0191] The tongue position height is used to evaluate the vertical position of the tongue in the mouth, and the calculation formula is:
[0192] H = ∥p tongu_center - p mout_bottom ∥
[0193] where H is the tongue position height, p tongu_center is the coordinate of the center point of the tongue, p mout_bottom is the coordinate of the center point at the bottom of the mouth.
[0194] Calculation formula for the relative position of the mouth shape and the tongue:
[0195] The relative position of the mouth shape and the tongue is used to further evaluate the dynamic position relationship of the tongue in the mouth, and the calculation formula is:
[0196]
[0197] where Relative_Position is the relative position of the tongue and the opening / closing of the mouth shape, H is the tongue position height, and R is the mouth opening / closing ratio.
[0198] In this embodiment, the data fusion and model comparison unit performs weighted processing on audio and visual features, performs matching calculations with the standard pronunciation model, identifies the pronunciation problems of the practitioner, and locates the deviation content.
[0199] Multimodal data fusion:
[0200] Audio feature matching score F audio and visual feature matching score F visual are fused through weighted processing:
[0201] F combined = w1·F audio + w2·F visual
[0202] w1, w2: Weight parameters of audio and visual features, satisfying w1 + w2 = 1
[0203] Deviation identification:
[0204] Deviation calculation formula:
[0205] Δ error =|F combined -F standard |
[0206] F standard : The eigenvalue of the standard pronunciation model.
[0207] Δ error : Feature deviation amount.
[0208] Problem location:
[0209] According to the deviation value Δ error Determine the pronunciation problem of the practitioner:
[0210] If Δ error >Threshold, output the specific problem type (such as insufficient mouth opening and closing, low tongue position).
[0211] Working principle: The device of the present invention is based on a notebook folding structure, and combines voice collection, visual detection, entertainment interaction and multimodal feedback functions to achieve accurate training and interesting learning of children's English pronunciation.
[0212] When in use, the practitioner turns on the device and selects a learning mode (word, sentence or free pronunciation) by operating the pen 26. During the practice, the carbon steel spring 7 is rotated out from the inside of the storage slot 1 5 by adjusting the seat 6, and then the practitioner adjusts the direction and position of the microphone 8 according to the height of his or her mouth position, accurately collecting the practitioner's pronunciation data, including pitch, tone, duration, etc., while removing environmental noise. At the same time, the camera 11 is pulled out from the inside of the storage slot 2 9 through the folding frame 10, and then the direction and position of the camera 11 are adjusted according to the practitioner's sitting position, so that it is aimed at the practitioner's neck and mouth, and the practitioner's mouth shape, tongue position and Adam's apple dynamics are captured in real time.
[0213] These data are transmitted to the analysis module of the system. The system compares the user's pronunciation data with the standard pronunciation model through speech spectrum analysis and computer vision algorithms, and identifies the type of pronunciation problems in real time, such as inaccurate tongue position or substandard mouth shape.
[0214] The system analysis results trigger a multi-modal feedback mechanism to guide the practitioner to adjust their pronunciation. When the system detects pronunciation, the entertainment interaction module uses the magnetic coil 15 to drive the ejector rod 16 to vibrate through electromagnetic drive, providing tactile cues: a light frequency vibration indicates correct pronunciation, a medium frequency vibration indicates a slight deviation, and a high frequency vibration alerts a significant error. At the same time, the voice feedback module plays fun prompts, such as "Raise your tongue a little higher, like licking a candy!" or "Open your mouth wider, like a big tiger!" The dynamic display screen shows a mouth shape comparison diagram and adjustment animation in real time, intuitively guiding the practitioner to adjust their pronunciation posture. At the same time, the ejector rod 16 triggers the attracting member 19 to rise from the embedding groove 13, simulating a cartoonish scenario: the motor 21 drives the moving member 24 to grab the attracting member 19, and at the same time plays a voice prompt, such as "The little rabbit is about to steal the radish, quickly stop it with the correct pronunciation!" If the pronunciation is correct, the attracting member 19 falls back, and the system plays a reward animation or voice to encourage.
[0215] The entire system forms a closed loop from data collection to feedback, interaction, and summary. The device records the pronunciation accuracy rate and problem types of each practice and generates a learning report to show the user's learning progress and trends, which not only helps parents understand the learning effect of children but also encourages children to continuously improve their pronunciation.
[0216] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An English pronunciation practice assistance device, including a storage board (1), characterized in that, A sound collection component and a camera component are provided on the surface of the storage board (1). An entertainment area (12) is provided on the upper surface of the storage board (1), and an entertainment mechanism is provided in the entertainment area (12). The entertainment mechanism includes a vibration component and a moving component; The storage board (1) is connected to a display board (3) through a rotating shaft (2). A display screen (4) is installed on the outer wall of the display board (3), and the display screen (4) is electrically connected to a voice-assisted learning system.
2. The auxiliary device for English pronunciation practice according to claim 1, characterized in that The sound collection component includes a storage groove one (5). The storage groove one (5) is opened on the upper surface of the storage board (1). An adjusting seat (6) rotates on the inner wall of the storage groove one (5). One end of a carbon steel spring (7) is installed at the end of the adjusting seat (6), and the other end of the carbon steel spring (7) is installed with a sound collection cylinder (8).
3. An English pronunciation practice assistance device according to claim 1, characterized in that, The camera component includes a storage groove two (9). The storage groove two (9) is opened on the upper surface of the storage board (1). One end of a folding frame (10) rotates on the inner wall of the storage groove two (9), and the other end of the folding frame (10) rotates with a camera (11).
4. An English pronunciation practice assistance device according to claim 1, characterized in that, The vibration component includes multiple groups of buried grooves (13). Each group of buried grooves (13) is opened at the bottom of the entertainment area (12). Multiple groups of guiding grooves (14) are opened inside the storage board (1). Each guiding groove (14) communicates with the buried groove (13). A limiting convex block (18) is installed at the connection of the buried groove (13) and the guiding groove (14). A magnetic ring (15) is arranged in the guiding groove (14), a push rod (16) is arranged in the magnetic ring (15), an attracting piece (19) is installed at the top of the push rod (16), a clamping ring (17) is installed at the bottom of the push rod (16), and the outer diameter of the clamping ring (17) is larger than the inner diameter of the limiting convex block (18).
5. An English pronunciation practice assistance device according to claim 1, characterized in that, The moving component includes multiple installation grooves (20). Each installation groove (20) is opened on the surface of the entertainment area (12). A motor (21) is installed at one end of the inner wall of the installation groove (20). A reciprocating lead screw (22) is installed at the driving end of the motor (21). A moving seat (23) is threaded on the outer wall of the reciprocating lead screw (22), and a moving piece (24) is installed at the top of the moving seat (23).
6. The auxiliary device for English pronunciation practice according to claim 5, characterized in that, Multiple groups of the buried grooves (13) and multiple of the installation grooves (20) are arranged alternately.
7. An English pronunciation practice assisting device according to claim 1, characterized in that, A storage groove three (25) is opened at the top of the storage board (1), and an operating pen (26) is arranged in the storage groove three (25).
8. An English pronunciation practice assisting device according to claim 1, characterized in that, The voice-assisted learning system includes: A data collection module, which is used to collect the voice data and visual data of the practitioner in real time as the basis for pronunciation analysis; A data analysis module, which is used to compare the collected data with the standard pronunciation model through voice spectrum analysis and visual algorithms to identify the types of pronunciation problems; A vibration feedback module, which is used to provide tactile feedback according to the system analysis result, and prompt whether the pronunciation of the practitioner is correct through vibrations of different frequencies; A voice feedback module, which is used to provide voice prompts according to the pronunciation problems to guide the practitioner to adjust the pronunciation posture; A visualization feedback module, which is used to dynamically display the lip comparison diagram, adjustment animation and reward animation through the display screen (4) to intuitively guide the practitioner to adjust the pronunciation; Entertainment interaction module, which is used to simulate a cartoonized scenario during the practice process and attract the attention of the practitioner through the dynamic actions of the attracting part (19) and the moving part (24); Learning record and summary module, which is used to record the pronunciation data and problem types of each practice, generate a learning report, and display the user's learning progress and trends.
9. An English pronunciation practice assisting device according to claim 8, characterized in that, The data analysis module includes: Audio processing unit, which is used to collect the pronunciation data of the practitioner, perform sampling, preprocessing, feature extraction and spectrum analysis, extract pitch, tone and duration features, and perform matching calculations with the standard pronunciation model; Visual processing unit, which is used to capture the dynamic movements of the practitioner's mouth shape, tongue position and Adam's apple, extract the key point coordinates and dynamic features, calculate the mouth opening and closing ratio and the tongue position, and judge the difference from the standard model; Data fusion and model comparison unit, which is used to perform weighted processing on audio and visual features, perform matching calculations with the standard pronunciation model, identify the pronunciation problems of the practitioner and locate the deviation content.
10. An English pronunciation practice assisting device according to claim 9, characterized in that, The specific calculation of the mouth opening and closing ratio and the tongue position in the visual processing unit includes calculating the mouth opening and closing ratio, the height of the tongue position and the relative position between the mouth shape and the tongue. The formulas are as follows: Formula for calculating the mouth opening and closing ratio, which is used to evaluate the degree of mouth opening and closing. It is calculated by the distance between the key points of the upper and lower lips and the key points of the corners of the mouth. Specifically: Among them, R is the mouth opening and closing ratio, p uppe_lip is the coordinate of the midpoint of the upper lip, p lower_lip is the coordinate of the midpoint of the lower lip, p left_corner is the coordinate of the left mouth corner, p right_corner is the coordinate of the right mouth corner, |·| is the Euclidean distance, representing the straight-line distance between two points; Formula for calculating the height of the tongue position, which is used to evaluate the vertical position of the tongue in the mouth. It is calculated by the distance between the key points of the tongue and the key points at the bottom of the mouth. Specifically: H = ∥p tongu_center -p mout_bottom ∥ Among them, H is the height of the tongue position, p tongu_center is the coordinate of the center point of the tongue, p mout_bottom is the coordinate of the center point of the bottom of the mouth; Formula for calculating the relative position between the mouth shape and the tongue, which is used to further evaluate the relative dynamics of the tongue position in the mouth. Specifically: Among them, Relative_Position is the relative position between the tongue and the mouth opening and closing, H is the height of the tongue position, and R is the mouth opening and closing ratio.
Citation Information
Cited By
English follow-up reading device for English teaching
CN120833693A