Teaching appliance for English teaching and scoring method
By designing an English teaching aid with a support frame that can be locked in a vertical posture and an extended handle, combined with end-to-end acoustic models and advanced processing technology, the problems of poor flexibility and noise interference of existing English teaching aids are solved, and accurate evaluation of English pronunciation is achieved.
Patent Information
- Application Number
- CN202510682104.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The support rod assembly and voiceless consonant lifting assembly of existing English teaching tools are difficult to store in the base, which is inconvenient for daily use and storage protection. The structure of the sound collection part also makes it difficult to extend the collection position, resulting in poor flexibility.
An English teaching tool was designed, which includes a support frame with a lockable vertical posture, an extendable handle and a sound collection component. Combining an end-to-end acoustic model and advanced processing technology, it accurately evaluates English pronunciation through preprocessing, feature extraction and scoring models.
It improves the flexibility of sound collection, effectively solves the problems of environmental noise interference and inconsistent feature dimensions, and achieves accurate evaluation of English pronunciation, which has high practical value and application prospects.
Smart Images

Figure CN120612852A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of teaching tools, and in particular relates to a teaching tool and a scoring method for English teaching. Background Art
[0002] English, as a second language, is widely used in various fields. Improving pronunciation is a key aspect of English learning. English has 28 consonants, with a rich contrast between voiced and unvoiced sounds. This is a significant advantage for Chinese speakers, as even if they struggle to pronounce voiced sounds, they can still distinguish between them. Speech assessment plays a key role in measuring English learners' oral proficiency. Traditionally, speech assessment relies primarily on specific audio acquisition equipment and standardized assessment processes. With technological advancements, the use of microphones has become increasingly widespread, enabling more flexible collection of English speech. Existing speech assessment technologies primarily rely on traditional hidden Markov modeling combined with goodness-of-pronunciation (GOP) algorithms, and on deep features derived from end-to-end acoustic models. However, these technologies face challenges when processing English speech collected by microphones, such as poor handling of ambient noise and an inability to fully utilize the characteristics of the microphones.
[0003] Speech evaluation technologies based on traditional hidden Markov modeling combined with GOP algorithms are susceptible to interference from ambient noise when processing audio collected by microphones, leading to inaccurate calculation of phoneme likelihood values and thus affecting scoring. Speech evaluation technologies based on deep features output by end-to-end acoustic models, while offering certain advantages in scoring, can also be affected by blank spaces in the microphone-collected audio and ambient noise.
[0004] There is an existing related patent, patent application number 202122122113.7, "A spoken language practice prompt device for English teaching", which has the characteristics of "this utility model is used to practice the pronunciation of consonants in English, distinguish between voiceless consonants and voiced consonants, and design corresponding detection devices according to the two pronunciation parts. According to the sound production principles of breathy sounds and glottal sounds, a patch-type vibration sensor for detecting breathy sounds and glottal sounds is designed." However, it still has some shortcomings, such as its support rod assembly and voiceless consonant lifting assembly are difficult to store in the base, which is not convenient for daily use and storage protection. The structure of the sound collection part is also difficult to extend the collection position, resulting in poor flexibility. Therefore, a teaching aid and scoring method for English teaching are proposed. Summary of the Invention
[0005] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.
[0006] In view of the following technical problems in the existing technology: the support rod assembly and the voiceless consonant lifting assembly are difficult to store in the base, which is inconvenient for daily use and storage protection; the structure of the sound collection part is also difficult to extend the collection position, resulting in poor flexibility.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions: a teaching aid for English teaching and a scoring method, comprising a body, a U-shaped frame is provided on the upper side of the back of the body, the U-shaped frame is rotatably connected to a cover plate, a top end of the body is recessed with a receiving slot, the receiving slot is screwed to a support frame, the support frame is recessed with a plug-in slot, a handle is movably plugged into the plug-in slot, a sound collecting component is provided on the handle, a sound collecting component is provided in the sound collecting component, a connecting wire 1 is connected to the sound collecting component, a rotating cavity is recessed on the inner bottom wall of the receiving cavity, the connecting wire 1 passes through the handle and the support frame and reaches the rotating cavity, a winding wheel is rotatably connected in the rotating cavity, a torsion spring is connected between the winding wheel and the rotating cavity, and the winding wheel is wound around the connecting wire 1;
[0008] A lifting plate is provided at the middle of one end of the cover away from the machine body.
[0009] As an optimal technical solution for a teaching tool and scoring method for English teaching, the support frame is movably connected with a locking pin, and the receiving groove is recessed with a locking hole, and the locking pin is plugged into the locking hole;
[0010] The locking pin is plugged into the locking hole to lock the support frame into a vertical position.
[0011] As an optimal technical solution for a teaching tool and scoring method for English teaching, a control ring is provided at one end of the locking pin, and a cylindrical magnet is provided at the middle of the locking pin;
[0012] The cylindrical magnet is adsorbed to the inside of the support frame and is difficult to separate.
[0013] As a preferred technical solution for a teaching aid and scoring method for English teaching, a movable folding frame is provided at one end of the body, and the movable folding frame includes a placement table 1, a placement table 2, a long cavity, and a connecting rod. A mounting seat is provided on one side of the body, and the mounting seat is rotatably connected to the placement table 1. The placement table 1 and the placement table 2 are both recessed with a long cavity. One end of the connecting rod extends into the long cavity of the mounting seat, and the other end of the connecting rod extends into the long cavity of the placement table 1.
[0014] The second placement platform is controlled to move away from the first placement platform, and the connecting rod slides in the long cavity to lengthen the movable folding frame, and then the movable folding frame is pulled to a horizontal posture, and then can be used to carry processing equipment.
[0015] As an optimal technical solution for a teaching tool and scoring method for English teaching, a display screen is provided on one side of the cover plate, a magnetic plate is provided on the other side of the display screen, a plurality of circular magnets are provided on the magnetic plate, and the circular magnets magnetically attract a magnetic sheet, a plastic sheet is provided on the front of the magnetic sheet, and English words or phonetic symbols are provided on the plastic sheet;
[0016] The display screen is connected to a second connecting cable, which is provided with an HDMI connector for connecting to a processing device for interaction.
[0017] As a preferred technical solution for an English teaching tool and scoring method, a sound collecting element is provided with a sound receiving tube, a through-groove is provided on the top of the sound receiving tube, four through-grooves are evenly provided on the circumference of the sound receiving tube, an elastic film is provided on the through-grooves, and the sound receiving element is located on the inner bottom wall of the sound receiving tube;
[0018] The elastic film allows sound vibrations to be transmitted to the sound receiver, and the sound is transmitted to the sound receiving element through the air in the sound receiver. The elastic film can prevent liquid from entering the sound receiver.
[0019] As a preferred technical solution for a teaching tool and scoring method for English teaching, an annular matching groove is recessed in the plug-in slot, and a mounting round seat is provided between the sound collecting component and the handle, and the mounting round seat is plugged into the annular matching groove;
[0020] The mounting round seat is provided with an annular magnetic seat, and the annular magnetic seat is magnetically attracted to the inner wall of the annular matching groove;
[0021] The magnetic seat and the inner wall of the annular matching groove are magnetically attracted to each other so that the sound collecting component and the handle are firmly locked on the support frame.
[0022] As an optimal technical solution for a teaching aid and scoring method for English teaching, a processing device is provided after the movable folding frame is unfolded, and the processing device is connected to the connecting line 2;
[0023] S1: The processing device uses the English speech audio collected by the sound receiver: The sound receiver collects English speech audio data, which may contain various environmental noises and background sounds;
[0024] S2: Preprocessing and feature extraction of audio:
[0025] S2.1: Preprocessing: This includes pre-emphasis, windowing, and discrete Fourier transform. Pre-emphasis can compensate for the loss of high-frequency components.
[0026] S2.2: Windowing divides a longer audio period into shorter periods with statistical characteristics, and discrete Fourier transform converts the time-domain audio data into frequency-domain data for subsequent analysis.
[0027] S2.3: Feature extraction: Extract frequency domain features Fbank from the preprocessed audio. Fbank features can retain acoustic feature information to the greatest extent and are suitable for subsequent feature extraction and scoring operations.
[0028] S3: Input the audio features into the trained scoring model:
[0029] S3.1: Scoring model structure: The scoring model consists of an end-to-end acoustic model, a noise filtering module, a self-attention mechanism module, a long short-term memory network, and multiple linear layers connected in sequence;
[0030] S3.2: End-to-end acoustic model: It consists of an encoder and a decoder connected in sequence. The encoder uses a Conformer encoder to obtain deep features from the frequency domain Fbank. The decoder uses a connection temporal classification model (CTC) to obtain the index value corresponding to the non-blank frame.
[0031] S3.3: Noise filtering module: This module filters the deep features of the end-to-end acoustic model based on the index values corresponding to the non-blank frames obtained by the decoder, removes the deep features corresponding to the blank and noisy frames, and reduces the interference of ambient noise and blank audio on the scoring.
[0032] S3.4: Self-attention mechanism module: assigns different weights to the filtered non-blank depth features, so that the model pays more attention to the parts that have a greater impact on the final score;
[0033] S3.5: Long Short-Term Memory Network: Input the features into the Long Short-Term Memory Network and obtain only the cell states of the Long Short-Term Memory Network to unify the dimensions of audio features of different lengths.
[0034] S3.6: Multi-layer linear layer: The LSTM features are fed into the multi-layer linear layer. After linear transformation, activation function, and normalization, the layer outputs a score between 0 and 1.
[0035] S4: Give evaluation based on the scoring results: In actual use, the output score between 0 and 1 is proportionally magnified according to the actual scoring system and rounded to obtain the final rating. At the same time, specific evaluation suggestions are given based on the scoring results, such as the strengths and weaknesses in pronunciation accuracy, fluency, intonation, etc.
[0036] In a second aspect, the present invention provides a system for evaluating English utterances collected by a microphone.
[0037] S4.1: Acquisition module: configured to acquire English voice audio collected by the pickup;
[0038] S4.2: Feature extraction module: configured to preprocess and extract features from the audio to obtain English speech audio features;
[0039] S4.3: Scoring module: configured to input the audio features into the trained scoring model and output a scoring result of the English utterance;
[0040] S4.4: Evaluation module: configured to give specific evaluation suggestions based on the scoring results.
[0041] The English teaching aid and scoring method of the present invention have the following beneficial effects: through the use of the support frame, the sound collecting member and the handle, the support frame can be locked in a vertical position, and the sound collecting member and the handle can also be separated from the support frame. By pulling the connecting line and continuously separating from the reel, the sound collecting member can be extended and used, making the sound collection more flexible.
[0042] For the collected English pronunciation evaluation problem, combining end-to-end acoustic models and advanced processing technology, it effectively solves problems such as environmental noise interference and inconsistent feature dimensions, and realizes accurate evaluation of English pronunciation, which has high practical value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:
[0044] Figure 1 It is a schematic diagram of the overall structure of the present invention;
[0045] Figure 2 Schematic diagram of the internal structure of the sound collection component of the present invention;
[0046] Figure 3 It is a schematic diagram of the three-dimensional structure of the microphone of the present invention;
[0047] Figure 4 This is a schematic structural diagram of the locking pin of the present invention;
[0048] Figure 5 This is a schematic diagram of the connection relationship between the locking pin and the support frame of the present invention;
[0049] Figure 6 For the present invention Figure 1 A schematic diagram of the partially enlarged structure of part A;
[0050] Figure 7 For the present invention Figure 1 Schematic diagram of the locally enlarged structure of part B.
[0051] Figure numerals: body -1, support frame -2, connecting line 1 -3, accommodating slot -4, handle -5, reel -6, rotating cavity -7, plug-in slot -8, display screen -9, locking pin -10, lifting plate -11, magnetic plate -12, magnetic sheet -13, cover plate -14, sound collecting component -15, connecting line 2 -16, adapter -17, mounting seat -18, placing table 1 -19, placing table 2 -20, long cavity -21, connecting rod -22, folding platform -23, connecting sheet -24, mounting round seat -25, annular magnetic seat -26, sound receiving component -27, sound receiving tube -28, elastic film -29, cylindrical magnet -30, annular matching groove -31. DETAILED DESCRIPTION
[0052] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0053] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0054] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0055] Furthermore, the present invention is described in detail with reference to schematic diagrams. For ease of illustration, when describing the embodiments of the present invention, cross-sectional views illustrating device structures may be partially enlarged and not to scale. Furthermore, the schematic diagrams are merely illustrative and should not limit the scope of protection of the present invention. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.
[0056] Example 1: Figures 1 to 7As shown, the present invention proposes a teaching tool for English teaching and a scoring method, comprising a body 1, a U-shaped frame is provided on the upper side of the back of the body 1, the U-shaped frame is rotatably connected to a cover plate 14, a top of the body 1 is recessed with a receiving slot 4, the receiving slot 4 is screwed to a support frame 2, a plug-in slot 8 is recessed on the support frame 2, a handle 5 is movably plugged into the plug-in slot 8, a sound collecting component 15 is provided on the handle 5, a sound receiving component 27 is provided in the sound collecting component 15, and a connecting wire 3 is connected to the connecting wire 27, a rotating cavity 7 is recessed on the inner bottom wall of the receiving groove 4, the connecting wire 3 passes through the handle 5 and the support frame 2 and arrives in the rotating cavity 7, a winding wheel 6 is rotatably connected in the rotating cavity 7, a torsion spring is connected between the winding wheel 6 and the rotating cavity 7, and the winding wheel 6 is wound around the connecting wire 3;
[0057] A lifting plate 11 is provided at the middle portion of the cover plate 14 away from the end of the body 1 .
[0058] The support frame 2 is movably connected with a locking pin 10, and the receiving groove 4 is provided with a locking hole, and the locking pin 10 is plugged into the locking hole;
[0059] The locking pin 10 is inserted into the locking hole to lock the support frame 2 in a vertical position.
[0060] A control ring is provided at one end of the locking pin 10, and a cylindrical magnet 30 is provided in the middle of the locking pin 10;
[0061] The cylindrical magnet 30 is adsorbed to the inside of the support frame 2 and is difficult to separate.
[0062] A movable folding frame is provided at one end of the machine body 1, and the movable folding frame includes a placing table 19, a placing table 20, a long cavity 21, and a connecting rod 22. A mounting seat 18 is provided on one side of the machine body 1, and the mounting seat 18 is rotatably connected to the placing table 19. The placing table 19 and the placing table 20 are both recessed with a long cavity 21. One end of the connecting rod 22 extends into the long cavity 21 of the mounting seat 18, and the other end of the connecting rod 22 extends into the long cavity 21 of the placing table 19.
[0063] The second placement platform 20 is controlled to move away from the first placement platform 19, and the connecting rod 22 slides in the long cavity 21 to lengthen the movable folding frame, and then the movable folding frame is pulled to a horizontal posture, and then can be used to carry processing equipment.
[0064] A display screen 9 is provided on one side of the cover plate 14, and a magnetic plate 12 is provided on the other side of the display screen 9. A plurality of circular magnets are provided on the magnetic plate 12, and the circular magnets magnetically attract a magnetic sheet 13. A plastic sheet is provided on the front of the magnetic sheet 13, and the plastic sheet is provided with English words or phonetic symbols;
[0065] The display screen 9 is connected to the second connection line 16, and the second connection line 16 is provided with an HDMI connector, which can be used to connect to a processing device for interaction.
[0066] A sound collecting member 15 is provided with a sound receiving tube 28. A through groove is provided at the top of the sound receiving tube 28. Four through grooves are evenly provided on the circumference of the sound receiving tube 28. An elastic film 29 is provided on the through groove. The sound receiving member 27 is located on the inner bottom wall of the sound receiving tube 28.
[0067] The elastic film 29 allows sound vibration to be transmitted to the sound receiving tube 28 , and the sound is transmitted to the sound receiving element 27 through the air in the sound receiving tube 28 . The elastic film 29 can prevent liquid from entering the sound receiving tube 28 .
[0068] An annular matching groove 31 is concavely provided in the plug-in slot 8, and a mounting round seat 25 is provided between the sound collecting member 15 and the handle 5, and the mounting round seat 25 is plugged into the annular matching groove 31;
[0069] The mounting round seat 25 is provided with an annular magnetic seat 26, and the annular magnetic seat 26 is magnetically attracted to the inner wall of the annular matching groove 31;
[0070] The magnetic base 26 and the inner wall of the annular matching groove 31 are magnetically attracted to each other so that the sound collecting component 15 and the handle 5 are firmly locked on the support frame 2 .
[0071] After the movable folding frame is unfolded, a processing device is provided, and the processing device is connected to the second connecting line;
[0072] Scoring Methodology:
[0073] S1: The processing device uses the English speech audio collected by the sound receiver: The sound receiver collects English speech audio data, which may contain various environmental noises and background sounds;
[0074] S2: Preprocessing and feature extraction of audio:
[0075] S2.1: Preprocessing: This includes pre-emphasis, windowing, and discrete Fourier transform. Pre-emphasis can compensate for the loss of high-frequency components.
[0076] S2.2: Windowing divides a longer audio period into shorter periods with statistical characteristics, and discrete Fourier transform converts the time-domain audio data into frequency-domain data for subsequent analysis.
[0077] S2.3: Feature extraction: Extract frequency domain features Fbank from the preprocessed audio. Fbank features can retain acoustic feature information to the greatest extent and are suitable for subsequent feature extraction and scoring operations.
[0078] S3: Input the audio features into the trained scoring model:
[0079] S3.1: Scoring model structure: The scoring model consists of an end-to-end acoustic model, a noise filtering module, a self-attention mechanism module, a long short-term memory network, and multiple linear layers connected in sequence;
[0080] S3.2: End-to-end acoustic model: consists of an encoder and a decoder connected in sequence; the encoder uses a Conformer encoder to obtain deep features from the frequency domain Fbank;
[0081] The decoder uses the connection temporal classification model CTC to obtain the index value corresponding to the non-blank frame;
[0082] S3.3: Noise filtering module: This module filters the deep features of the end-to-end acoustic model based on the index values corresponding to the non-blank frames obtained by the decoder, removes the deep features corresponding to the blank and noisy frames, and reduces the interference of ambient noise and blank audio on the scoring.
[0083] S3.4: Self-attention mechanism module: assigns different weights to the filtered non-blank depth features, so that the model pays more attention to the parts that have a greater impact on the final score;
[0084] S3.5: Long Short-Term Memory Network: Input the features into the Long Short-Term Memory Network and obtain only the cell states of the Long Short-Term Memory Network to unify the dimensions of audio features of different lengths.
[0085] S3.6: Multi-layer linear layer: The LSTM features are fed into the multi-layer linear layer. After linear transformation, activation function, and normalization, the layer outputs a score between 0 and 1.
[0086] S4: Give evaluation based on the scoring results: In actual use, the output score between 0 and 1 is proportionally magnified according to the actual scoring system and rounded to obtain the final rating. At the same time, specific evaluation suggestions are given based on the scoring results, such as the strengths and weaknesses in pronunciation accuracy, fluency, intonation, etc.
[0087] In a second aspect, the present invention provides a system for evaluating English utterances collected by a microphone.
[0088] S4.1: Acquisition module: configured to acquire English voice audio collected by the pickup;
[0089] S4.2: Feature extraction module: configured to preprocess and extract features from the audio to obtain English speech audio features;
[0090] S4.3: Scoring module: configured to input the audio features into the trained scoring model and output a scoring result of the English utterance;
[0091] S4.4: Evaluation module: configured to give specific evaluation suggestions based on the scoring results.
[0092] The processing device includes a computer, and the display screen 9 is connected to the second connecting line 16, and the second connecting line 16 is provided with an HDMI connector.
[0093] A mouse platform is provided on one side of the second placement platform 20. The mouse platform comprises a folding platform 23 and a connecting piece 24. The second placement platform 20 is connected to the folding platform 23 via the connecting piece 24. The folding platform 23 is L-shaped and the connecting piece 24 is pivotally connected to the folding platform 23. The connecting piece 24 is also pivotally connected to the second placement platform 20. A sound receiver 27 comprises a pickup or microphone. When the sound receiver 27 is a pickup, the model is DS-65VA300W. The pickup is directly connected to the second connection line 16 via the first connection line 3. The second connection line 16 is provided with a USB connector, which is connected to the processing device.
[0094] A channel is provided in the support frame 2 for the connecting line 13 to pass through.
[0095] The cover plate 14 is controlled to flip, so that the cover plate 14 can be separated from the upper surface of the body 1, or even vertically, so that the display screen 9 can be displayed and used. The support frame 2 is swung out from the receiving groove 4, and the locking pin 10 is inserted into the support frame 2 and the receiving groove 4. The cylindrical magnet 30 makes it difficult for the locking pin 10 to separate from the support frame 2, so that the support frame 2 is locked in the vertical position. The sound collecting member 15 and the handle 5 can be separated from the support frame 2. The connecting wire 3 is pulled to continuously separate from the reel 6, so that the sound collecting member 15 can be extended and used, making it more flexible to collect sound.
[0096] Control the placement table 2 20 away from the placement table 1 19, the connecting rod 22 slides in the long cavity 21 to make the movable folding frame longer, and then pull the movable folding frame to a horizontal posture, and then carry the processing equipment. The space between the placement table 20 and the placement table 1 19 is conducive to the suspended heat dissipation of the processing equipment. Control the folding platform 23 to be horizontal. A rubber layer is provided on the folding platform 23 for mouse movement. The USB or HDMI connector on the connecting line 2 16 is connected to the processing equipment for data transmission and sound collection.
[0097] This embodiment provides a method for evaluating English utterances collected by a microphone:
[0098] Get the English audio collected by the pickup:
[0099] Use a microphone to collect audio data of English pronunciation to ensure the integrity and accuracy of the audio data.
[0100] Preprocess and extract features from audio:
[0101] Preprocessing:
[0102] Pre-emphasis processing: As the frequency of sound increases, the medium loses more sound energy during propagation. To protect sound information, pre-emphasis is used to compensate for the loss of high-frequency components. The calculation process is shown in formula (1): y(n) = x(n) - μx(n-1); where x(n) represents the nth sampling point of the audio data, μ ranges from (0.95 to 0.99), and y(n) represents the data obtained by pre-emphasis of the nth sampling point of the audio data.
[0103] Windowing: Audio changes continuously over a long period of time and has no statistical characteristics, but audio over a short period of time can be considered unchanged, thus having statistical characteristics. The windowing operation uses a fixed window length, takes a fixed length of speech each time, and performs a discrete Fourier transform on the audio within the window. The above process is repeated in sequence according to a fixed window shift until the entire audio is windowed. The windowing calculation process is shown in formula (2): ω(n) = x(n)·h(n), where x(n) is the nth sampling point in the window, h(n) is the weight corresponding to x(n), and ω(n) represents the data obtained by windowing the (nth) sampling point in the window.
[0104] Discrete Fourier transform: Perform discrete Fourier transform on the audio within each window after the windowing operation, convert the waveform into a spectrum, and convert the data within the window from the time domain to the frequency domain. The DFT calculation process is shown in formula (3).
[0105]
[0106] Where x(n) is the time domain signal, X(k) is the frequency domain signal, N is the number of sampling points in the window, k is the frequency, n is the nth sampling point in the window, and j represents the imaginary unit.
[0107] Feature extraction:
[0108] Frequency domain features Fbank are extracted from the preprocessed audio. Because the human ear's sensitivity to sound is uneven, a Mel scale is used to linearly map frequency to ear sensitivity. A Mel filter bank is used to map the spectrum to the Mel frequency scale. The result obtained through the Mel filter bank is the FBank feature. FBank features have high correlation and can maximize the preservation of acoustic feature information, making them suitable for subsequent feature extraction in end-to-end acoustic models.
[0109] Feed the audio features into the trained scoring model
[0110] Scoring model structure
[0111] The scoring model consists of an end-to-end acoustic model, a noise filtering module, a self-attention mechanism module, a long short-term memory network, and a multi-layer linear layer connected in sequence.
[0112] End-to-end acoustic model
[0113] Encoder: The Conformer encoder is used. Its network structure consists of a convolutional downsampling layer (ConvolutionSubsampling), a linear layer, a random dropout layer (Dropout), and N Conformer network modules (ConformerBlock) connected in sequence (N is a positive integer). The Conformer encoder is used to obtain deep features from the frequency domain Fbank.
[0114] Decoder: The Connected Temporal Classification (CTC) model is used to obtain the index values corresponding to non-blank frames. The CTC decoder consists of the log_softmax module, which performs the softmax activation function and logarithm, the topk algorithm module, which returns the maximum element, and the nonzero algorithm module, which returns nonzero elements.
[0115] Noise filtering module:
[0116] According to the index values corresponding to the non-blank frames obtained by the decoder, the depth features of the end-to-end acoustic model are filtered, and only the depth features corresponding to the frames with index values in the list are retained, and the depth features corresponding to the frames with index values not in the list are deleted, and the depth features corresponding to blank frames and noise frames are eliminated.
[0117] Self-attention mechanism module:
[0118] The non-blank deep features obtained by the filtering operation are input into the self-attention mechanism module, which assigns different weights to them according to their impact on subsequent scoring.
[0119] Long Short-Term Memory Network:
[0120] The features are input into the long short-term memory network, and only the cell state of the long short-term memory network is obtained to achieve dimensional unification of audio features of different lengths.
[0121] Multiple linear layers:
[0122] The multi-layer linear network structure consists of a first linear layer, a first activation function layer, a normalization layer, a second linear layer, and a second activation function layer, all connected in sequence. The normalized features of the long short-term memory network are input into the multi-layer linear layer. After linear transformation, Sigmoid activation function, and normalization, a score between 0 and 1 is output.
[0123] Give evaluation based on the rating results
[0124] In actual use, the output score between 0 and 1 is scaled up according to the actual scale and rounded to the nearest integer to obtain the final rating. At the same time, specific evaluation suggestions are given based on the scoring results, such as strengths and weaknesses in pronunciation accuracy, fluency, intonation, etc.
[0125] Example 2: Figures 1 to 7 As shown:
[0126] This embodiment provides a system for evaluating English utterances collected by a microphone.
[0127] The acquisition module is configured to acquire the English audio collected by the pickup.
[0128] Feature extraction module: configured to preprocess and extract features from the audio to obtain English speech audio features.
[0129] Scoring module: is configured to input audio features into the trained scoring model and output the scoring results of English utterances.
[0130] Evaluation module: configured to give specific evaluation suggestions based on the scoring results.
[0131] Aiming at the problem of evaluating English speech collected by a microphone, the present invention combines an end-to-end acoustic model with advanced processing technology to effectively solve problems such as environmental noise interference and inconsistent feature dimensions, thereby achieving accurate evaluation of English speech, and having high practical value and application prospects.
[0132] It will be appreciated that in the development of any actual embodiment, as in any engineering or design project, numerous implementation-specific decisions may be made. Such a development effort may be complex and time-consuming, but will, for those of ordinary skill having the benefit of this disclosure, be a routine undertaking of design, fabrication, and production without undue experimentation.
[0133] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A teaching aid and scoring method for English teaching, characterized by: The invention comprises a body (1), a U-shaped frame is provided on the upper side of the back of the body (1), the U-shaped frame is rotatably connected to a cover plate (14), a top end of the body (1) is concavely provided with a receiving groove (4), the receiving groove (4) is screwed to a support frame (2), a plug-in groove (8) is concavely provided on the support frame (2), a handle (5) is movably plugged into the plug-in groove (8), a sound collecting component (15) is provided on the handle (5), a sound receiving component (27) is provided in the sound collecting component (15), the sound receiving component (27) is connected to a connecting line (3), a rotating cavity (7) is concavely provided on the inner bottom wall of the receiving groove (4), the connecting line (3) passes through the handle (5) and the support frame (2) and arrives in the rotating cavity (7), a reel (6) is rotatably connected in the rotating cavity (7), a torsion spring is connected between the reel (6) and the rotating cavity (7), and the reel (6) is wound around the connecting line (3); A lifting plate (11) is provided at the middle of one end of the cover plate (14) away from the machine body (1).
2. The English teaching tool and scoring method according to claim 1, wherein: The support frame (2) is movably connected with a locking pin (10), and a locking hole is concavely provided in the receiving groove (4), and the locking pin (10) is connected with the locking hole.
3. The English teaching tool and scoring method according to claim 2, wherein: A control ring is provided at one end of the locking pin (10), and a cylindrical magnet (30) is provided at the middle of the locking pin (10).
4. The English teaching tool and scoring method according to claim 1, wherein: A movable folding frame is provided at one end of the machine body (1), and the movable folding frame includes a placing table (19), a placing table (20), a long cavity (21), and a connecting rod (22). A mounting seat (18) is provided on one side of the machine body (1), and the mounting seat (18) is rotatably connected to the placing table (19). The placing table (19) and the placing table (20) are both recessed with a long cavity (21). One end of the connecting rod (22) extends into the long cavity (21) of the mounting seat (18), and the other end of the connecting rod (22) extends into the long cavity (21) of the placing table (19).
5. The English teaching tool and scoring method according to claim 1, wherein: A display screen (9) is provided on one side of the cover plate (14), and a magnetic plate (12) is provided on the other side of the display screen (9). A plurality of circular magnets are provided on the magnetic plate (12), and the circular magnets magnetically attract a magnetic sheet (13). A plastic sheet is provided on the front of the magnetic sheet (13), and English words or phonetic symbols are provided on the plastic sheet.
6. The English teaching tool and scoring method according to claim 1, wherein: A sound collecting part (15) is provided with a sound receiving tube (28), a top of the sound receiving tube (28) is provided with a through groove, four through grooves are evenly provided on the circumference of the sound receiving tube (28), an elastic film (29) is provided on the through groove, and the sound receiving part (27) is located on the inner bottom wall of the sound receiving tube (28).
7. The English teaching tool and scoring method according to claim 6, characterized in that: An annular matching groove (31) is concavely provided in the plug-in groove (8), and a mounting round seat (25) is provided between the sound collecting member (15) and the handle (5), and the mounting round seat (25) is plugged into the annular matching groove (31); The mounting round seat (25) is provided with an annular magnetic seat (26), and the annular magnetic seat (26) is magnetically attracted to the inner wall of the annular matching groove (31).
8. The English teaching tool and scoring method according to claim 1, wherein: After the movable folding frame is unfolded, a processing device is provided, and the processing device is connected to the second connecting line; S1: English audio data collected by the audio receiver: The audio receiver collects English audio data, which may contain various environmental noises and background sounds; S2: Preprocessing and feature extraction of audio: S2.1: Preprocessing: This includes pre-emphasis, windowing, and discrete Fourier transform. Pre-emphasis can compensate for the loss of high-frequency components. S2.2: Windowing divides a longer audio period into shorter periods with statistical characteristics, and discrete Fourier transform converts the time-domain audio data into frequency-domain data for subsequent analysis. S2.3: Feature extraction: Extract frequency domain features Fbank from the preprocessed audio. Fbank features can retain acoustic feature information to the greatest extent and are suitable for subsequent feature extraction and scoring operations. S3: Input the audio features into the trained scoring model: S3.1: Scoring model structure: The scoring model consists of an end-to-end acoustic model, a noise filtering module, a self-attention mechanism module, a long short-term memory network, and multiple linear layers connected in sequence; S3.2: End-to-end acoustic model: consists of an encoder and a decoder connected in sequence; the encoder uses a Conformer encoder to obtain deep features from the frequency domain Fbank; The decoder uses the connection temporal classification model CTC to obtain the index value corresponding to the non-blank frame; S3.3: Noise filtering module: This module filters the deep features of the end-to-end acoustic model based on the index values corresponding to the non-blank frames obtained by the decoder, removes the deep features corresponding to the blank and noisy frames, and reduces the interference of ambient noise and blank audio on the scoring. S3.4: Self-attention mechanism module: assigns different weights to the filtered non-blank depth features, so that the model pays more attention to the parts that have a greater impact on the final score; S3.5: Long Short-Term Memory Network: Input the features into the Long Short-Term Memory Network and obtain only the cell states of the Long Short-Term Memory Network to unify the dimensions of audio features of different lengths. S3.6: Multi-layer linear layer: The LSTM features are fed into the multi-layer linear layer. After linear transformation, activation function, and normalization, the layer outputs a score between 0 and 1. S4: Give evaluation based on the scoring results: In actual use, the output score between 0 and 1 is proportionally magnified according to the actual scoring system and rounded to obtain the final rating. At the same time, specific evaluation suggestions are given based on the scoring results, such as the strengths and weaknesses in pronunciation accuracy, fluency, intonation, etc. In a second aspect, the present invention provides a system for evaluating English utterances collected by a microphone. S4.1: Acquisition module: configured to acquire English voice audio collected by the pickup; S4.2: Feature extraction module: configured to preprocess and extract features from the audio to obtain English speech audio features; S4.3: Scoring module: configured to input the audio features into the trained scoring model and output a scoring result of the English utterance; S4.4: Evaluation module: configured to give specific evaluation suggestions based on the scoring results.
Citation Information
Patent Citations
Oral practice prompting device for English teaching
CN215770184U