Deep learning-based intelligent interactive teaching system and method for spoken English emotion
Through the deep learning-based intelligent interactive teaching system for oral English, combined with environmental detection, audio and video data evaluation and lip feature optimization, the problem of insufficient accuracy of loss-break judgment in context coordinated pronunciation changes in the existing technology is solved, and higher judgment accuracy and multimodal learning experience are achieved.
Patent Information
- Application Number
- CN202510391421.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the prior art, there are a large number of context-coordinated pronunciation changes in long sentences input by users during the intelligent interaction of English speaking, such as continuous reading, weak reading, and loss of explosion, resulting in insufficient accuracy in judging the context-coordinated pronunciation changes.
Provides an intelligent interactive teaching system for spoken English based on deep learning, including environmental detection and evaluation module, threshold comparison module, audio and video collection and processing module, optimization module and suggestion output module. By evaluating the user-entered intelligent interactive audio data, combining lip feature data, the environment and lip feature are optimized to improve the accuracy of loss-break judgment.
It has achieved the accuracy of judging the loss of explosion in context collaborative pronunciation changes, enhanced the multimodal learning experience of the English oral emotional intelligent interactive teaching system based on deep learning, and improved the environmental applicability of the system.
Smart Images

Figure CN119905085B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech evaluation, and particularly to an intelligent interactive teaching system and method for English oral emotion based on deep learning. Background Art
[0002] With the in-depth development of information technology and the field of artificial intelligence, deep learning, as a powerful machine learning technology, has made breakthroughs in many aspects such as speech recognition, natural language processing, and sentiment analysis. The development of these technologies provides a solid foundation for the construction of an intelligent interactive teaching system for English oral language, enabling the system to understand and evaluate the oral expressions of learners, including speech accuracy, grammar correctness, and emotional expression.
[0003] In the prior art, an intelligent interactive teaching system for English oral language evaluates whether an English pronunciation loses a plosive sound by reflecting the frequency and amplitude of the sound through acoustic spectrum features.
[0004] For example, a method for improving the pronunciation quality in English teaching disclosed in the invention patent with the publication number: CN111916106B includes: obtaining the voice information input by the user; recognizing the voice information to obtain the characteristic parameters of the voice information, and transmitting the characteristic parameters to a voice evaluation model; the voice evaluation model is used to evaluate the voice information according to the characteristic parameters to obtain a voice evaluation result; when the voice evaluation result is standard, reminding the user that the pronunciation is standard through an output device; when the voice evaluation result is not standard, obtaining the voice content corresponding to the voice information, and transmitting the voice content to a standard voice model; the standard voice model is used to obtain standard voice information according to the voice content, and output the standard voice information through an output device; comparing the voice information with the standard voice information to obtain corresponding voice guidance information; and transmitting the voice guidance information to the user through an output device to assist the user in pronunciation training.
[0005] For example, a method, device, equipment and storage medium for evaluating plosive sounds disclosed in the invention patent with the publication number: CN113077822B includes: obtaining the English voice to be evaluated; sending the English voice into a decoding graph forced alignment for recognition, and the decoding graph includes a first pronunciation path with plosive sounds and a second pronunciation path without plosive sounds; if the second pronunciation path in the decoding graph is adopted during the recognition process, it is evaluated that the pronunciation of the English voice loses a plosive sound.
[0006] However, in the process of implementing the technical solutions of the present invention in the embodiments of the present application, it is found that the above technologies have at least the following technical problems:
[0007] In the prior art, during the intelligent interaction of spoken English, there are a large number of co-articulation changes in the long sentences input by users, such as liaison, weak reading, aspiration loss, etc. There is a problem of insufficient accuracy in judging aspiration loss in co-articulation changes in context. Summary of the Invention
[0008] By providing a spoken English emotion intelligent interaction teaching system and method based on deep learning, the embodiments of the present application solve the problem of insufficient accuracy in judging aspiration loss in co-articulation changes in context in the prior art, and achieve the effect of improving the accuracy of judging aspiration loss in co-articulation changes in context.
[0009] The embodiments of the present application provide a spoken English emotion intelligent interaction teaching system and method based on deep learning, including: an environment detection and evaluation module, a threshold comparison module, an audio-video collection and processing module, an optimization module, and a suggestion output module; the environment detection and evaluation module: used to collect and evaluate the spoken English intelligent interaction environment data to obtain a spoken English intelligent interaction environment evaluation coefficient, and transmit the spoken English intelligent interaction environment evaluation coefficient to the threshold comparison module; the threshold comparison module: used to compare the threshold range of the spoken English intelligent interaction environment evaluation coefficient obtained in the environment detection and evaluation module with the reference threshold of the image, and compare the threshold of the aspiration loss judgment evaluation coefficient of the spoken English intelligent interaction audio obtained in the audio-video collection and processing module; the audio-video collection and processing module: used to collect and evaluate the spoken English intelligent interaction audio data input by the user to obtain an aspiration loss judgment evaluation coefficient of the spoken English intelligent interaction audio, and transmit the aspiration loss judgment evaluation coefficient of the spoken English intelligent interaction audio to the suggestion output module; the optimization module: used to optimize the spoken English intelligent interaction environment according to the threshold range comparison result of the spoken English intelligent interaction environment evaluation coefficient. If the optimized spoken English intelligent interaction environment evaluation coefficient still fails to pass the threshold range comparison, send an instruction to the audio-video collection and processing module to synchronously collect the spoken English intelligent interaction lip feature data when collecting the spoken English intelligent interaction audio data, and optimize the spoken English intelligent interaction lip feature according to the threshold comparison result of the spoken English intelligent interaction environment evaluation coefficient; the suggestion output module: used to output text, audio, and visual spoken English error correction suggestions according to the aspiration loss judgment evaluation coefficient of the spoken English intelligent interaction audio.
[0010] Further, the threshold comparison module includes a threshold range comparison unit and a reference threshold comparison unit for the image; the threshold range comparison unit: is used to obtain the threshold range of the evaluation coefficient of the intelligent oral English interaction environment from the database, compare the evaluation coefficient of the intelligent oral English interaction environment with the threshold range of the evaluation coefficient of the intelligent oral English interaction environment. If the evaluation coefficient of the intelligent oral English interaction environment is within the threshold range of the evaluation coefficient of the intelligent oral English interaction environment, it sends an instruction to the audio-video collection and processing module to collect and evaluate the input intelligent oral English interaction audio data of the user, and obtain the explosion loss judgment evaluation coefficient of the intelligent oral English interaction audio. If the evaluation coefficient of the intelligent oral English interaction environment is greater than or equal to the maximum value of the threshold range of the evaluation coefficient of the intelligent oral English interaction environment, it sends an instruction to the optimization module to perform dynamic weight allocation for spatial positioning of the intelligent oral English interaction environment. If the evaluation coefficient of the intelligent oral English interaction environment is less than the minimum value of the threshold range of the evaluation coefficient of the intelligent oral English interaction environment, it sends an instruction to the optimization module to perform frequency domain signal processing on the intelligent oral English interaction environment. After the dynamic weight allocation for spatial positioning or the frequency domain signal processing, it sends an instruction to the environment detection and evaluation module to re-collect and evaluate the data of the intelligent oral English interaction environment, obtain the evaluation coefficient of the intelligent oral English interaction environment, and transmit the evaluation coefficient of the intelligent oral English interaction environment to the threshold comparison module. If the evaluation coefficient of the intelligent oral English interaction environment is still not within the threshold range of the evaluation coefficient of the intelligent oral English interaction environment at this time, it sends an instruction to the audio-video collection and processing module to synchronously collect the intelligent oral English interaction lip feature data when collecting the intelligent oral English interaction audio data, and perform combined analysis on the intelligent oral English interaction audio data and the optimized intelligent oral English interaction lip feature data to obtain the explosion loss judgment evaluation coefficient of the intelligent oral English interaction audio. If the evaluation coefficient of the intelligent oral English interaction environment is within the threshold range of the evaluation coefficient of the intelligent oral English interaction environment at this time, it sends an instruction to the audio-video collection and processing module to collect and evaluate the input intelligent oral English interaction audio data of the user, and obtain the explosion loss judgment evaluation coefficient of the intelligent oral English interaction audio;The reference threshold comparison unit of the image: It is used to obtain the sharpness reference threshold of the image in the intelligent oral English interaction from the database, obtain the sharpness of the image in the intelligent oral English interaction through the sensor, compare the sharpness of the image in the intelligent oral English interaction with the reference threshold of the sharpness of the image in the intelligent oral English interaction. If the sharpness of the image in the intelligent oral English interaction is less than or equal to the reference threshold of the sharpness of the image in the intelligent oral English interaction, it sends an instruction to the optimization module to optimize the lip feature data in the intelligent oral English interaction, and combines and analyzes the audio data in the intelligent oral English interaction with the optimized lip feature data in the intelligent oral English interaction to obtain the explosion loss judgment evaluation coefficient of the audio in the intelligent oral English interaction. If the sharpness of the image in the intelligent oral English interaction is greater than the reference threshold of the sharpness of the image in the intelligent oral English interaction, it sends an instruction to the audio and video collection and processing module to combine and analyze the audio data in the intelligent oral English interaction with the lip feature data in the intelligent oral English interaction to obtain the explosion loss judgment evaluation coefficient of the audio in the intelligent oral English interaction.;
[0011] Further, the specific process of collecting the intelligent oral English interaction environment data is as follows: Collect the intelligent oral English interaction environment data through the sensor, preprocess the intelligent oral English interaction environment data to obtain the preprocessed intelligent oral English interaction environment data, and perform preliminary noise reduction on the preprocessed intelligent oral English interaction environment data; The intelligent oral English interaction environment data includes the noise sound pressure level of the intelligent oral English interaction environment, the light intensity of the intelligent oral English interaction environment, and the color temperature of the intelligent oral English interaction environment.
[0012] Further, the specific process of collecting the intelligent oral English interaction audio data input by the user is as follows: Collect the intelligent oral English interaction audio data through the sensor, preprocess the intelligent oral English interaction audio data to obtain the preprocessed intelligent oral English interaction audio data; The intelligent oral English interaction audio data includes the phoneme combination coverage rate of the audio and the instantaneous energy peak value of the audio.
[0013] Further, the specific method for obtaining the evaluation coefficient of the intelligent oral English interaction environment is as follows: Obtain the weight factors of the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment from the database. Arrange the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment in a time series. Assign weights to the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment. Compare the standard values and the real-time values of the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment at each time series point. After averaging, the evaluation coefficient of the intelligent oral English interaction environment is obtained.
[0014] Further, the specific method for obtaining the evaluation coefficient of the aspiration judgment of the intelligent oral English interaction audio is as follows: When lip feature data needs to be combined, obtain the weight factors of the lip opening and closing distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value of the audio from the database. Assign weights to the deviation value of the lip opening and closing distance from the reference value of the lip opening and closing distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value of the audio. Sort the lip opening and closing distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value according to the number of phonemes. Combine and process the deviation value of the lip opening and closing distance from the reference value of the lip opening and closing distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value to obtain the evaluation coefficient of the aspiration judgment of the intelligent oral English interaction audio; When lip features do not need to be combined, obtain the weight factors of the phoneme combination coverage rate of the audio and the instantaneous energy peak value of the audio from the database. Assign weights to the phoneme combination coverage rate of the audio and the instantaneous energy peak value of the audio. Sort the phoneme combination coverage rate of the audio and the instantaneous energy peak value according to the number of phonemes. Combine and process the phoneme combination coverage rate of the audio and the instantaneous energy peak value to obtain the evaluation coefficient of the aspiration judgment of the intelligent oral English interaction audio.
[0015] Further, the specific optimization steps for optimizing the English spoken language intelligent interaction environment based on the comparison result of the threshold range of the English spoken language intelligent interaction environment evaluation coefficient are as follows: If the English spoken language intelligent interaction environment evaluation coefficient is greater than or equal to the maximum value of the threshold range of the English spoken language intelligent interaction environment evaluation coefficient, send an instruction to the optimization module to perform dynamic weight allocation for spatial positioning of the English spoken language intelligent interaction environment. If the English spoken language intelligent interaction environment evaluation coefficient is less than the minimum value of the threshold range of the English spoken language intelligent interaction environment evaluation coefficient, send an instruction to the optimization module to perform frequency domain signal processing on the English spoken language intelligent interaction environment.
[0016] Further, the specific steps for synchronously collecting English spoken language intelligent interaction lip feature data when collecting English spoken language intelligent interaction audio data are as follows: Obtain English spoken language intelligent interaction lip data through a depth sensor, extract English spoken language intelligent interaction lip feature data through a deep learning model, and use hardware triggering or software triggering to synchronize audio and video; The English spoken language intelligent interaction lip feature data includes the opening and closing distance of the lips; Hardware triggering includes using a synchronization signal generator to send a pulse signal to trigger the audio and video devices simultaneously; Software triggering includes achieving audio-visual stream timestamp alignment through the NDI protocol.
[0017] Further, the specific output method for outputting text, audio, and visual English spoken language error correction suggestions based on the aspiration judgment evaluation coefficient of the English spoken language intelligent interaction audio: Match predefined text, audio, and visual English spoken language error correction suggestions from the database according to the aspiration judgment evaluation coefficient of the English spoken language intelligent interaction audio. The text, audio, and visual English spoken language error correction suggestions include displaying spectrograms of the user's pronunciation and the standard pronunciation, playing back the user's English spoken language recording and highlighting the error parts, and pointing out the pronunciation error type and providing correct written pronunciation examples.
[0018] Furthermore, the deep learning-based intelligent interactive teaching method for spoken English is as follows: collect and evaluate the data of the intelligent interactive environment for spoken English to obtain the evaluation coefficient of the intelligent interactive environment for spoken English, and transmit the evaluation coefficient of the intelligent interactive environment for spoken English to the threshold comparison module; compare the evaluation coefficient of the intelligent interactive environment for spoken English obtained in the environment detection and evaluation module with the threshold range and the reference threshold of the image, and compare the evaluation coefficient of the explosion loss judgment of the intelligent interactive audio for spoken English obtained in the audio and video collection and processing module with the threshold; collect and evaluate the intelligent interactive audio data of spoken English input by the user to obtain the evaluation coefficient of the explosion loss judgment of the intelligent interactive audio for spoken English, and transmit the evaluation coefficient of the explosion loss judgment of the intelligent interactive audio for spoken English to the suggestion output module; optimize the intelligent interactive environment for spoken English according to the comparison result of the threshold range of the evaluation coefficient of the intelligent interactive environment for spoken English. If the evaluation coefficient of the optimized intelligent interactive environment for spoken English still fails to pass the threshold range comparison, send an instruction to the audio and video collection and processing module to synchronously collect the lip shape feature data of the intelligent interactive audio for spoken English when collecting the intelligent interactive audio data of spoken English, and optimize the lip shape feature of the intelligent interactive audio for spoken English according to the comparison result of the threshold of the evaluation coefficient of the intelligent interactive environment for spoken English; output text, audio and visual English spoken language error correction suggestions according to the evaluation coefficient of the explosion loss judgment of the intelligent interactive audio for spoken English.
[0019] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0020] 1. By evaluating the intelligent interactive audio data of spoken English input by the user, the evaluation coefficient of the explosion loss judgment of the intelligent interactive audio for spoken English is obtained, and thus text, audio and visual English spoken language error correction suggestions are output according to the evaluation coefficient of the explosion loss judgment of the intelligent interactive audio for spoken English, thereby achieving the effect of improving the accuracy of judging the explosion loss in the context collaborative pronunciation change, and effectively solving the problem of insufficient accuracy of judging the explosion loss in the context collaborative pronunciation change in the prior art.
[0021] 2. By synchronously collecting the lip shape feature data of the intelligent interactive audio for spoken English when collecting the intelligent interactive audio data of spoken English, the lip shape feature of the intelligent interactive audio for spoken English is optimized according to the comparison result of the threshold of the evaluation coefficient of the intelligent interactive environment for spoken English, thereby enhancing the multi-modal learning experience of the deep learning-based intelligent interactive teaching system for spoken English emotions, and effectively solving the problem of insufficient multi-modal learning experience of the deep learning-based intelligent interactive teaching system for spoken English emotions in the prior art.
[0022] 3. By comparing the evaluation coefficient of the intelligent English oral interaction environment obtained in the environment detection and evaluation module with the threshold range, the intelligent English oral interaction environment is optimized according to the comparison result of the threshold range of the evaluation coefficient of the intelligent English oral interaction environment, thereby improving the environmental applicability of the English oral emotion intelligent interaction teaching system based on deep learning and effectively solving the problem of insufficient environmental applicability of the English oral emotion intelligent interaction teaching system based on deep learning in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 FIG. 1 is a schematic structural diagram of an English oral emotion intelligent interaction teaching system based on deep learning provided by an embodiment of the present application;
[0024] Figure 2 FIG. 2 is a schematic structural diagram of a threshold comparison module in an English oral emotion intelligent interaction teaching system based on deep learning provided by an embodiment of the present application;
[0025] Figure 3 FIG. 3 is a schematic diagram of an evaluation coefficient of an intelligent English oral interaction environment in an English oral emotion intelligent interaction teaching system based on deep learning provided by an embodiment of the present application;
[0026] Figure 4 FIG. 4 is a schematic diagram of an English oral emotion intelligent interaction teaching method based on deep learning provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] By providing an English oral emotion intelligent interaction teaching system and method based on deep learning, the embodiment of the present application solves the problem of insufficient accuracy in judging elision in context co-articulation changes in the prior art. By evaluating the English oral intelligent interaction audio data input by the user, an elision judgment evaluation coefficient of the English oral intelligent interaction audio is obtained, and thus text, audio, and visual English oral error correction suggestions are output according to the elision judgment evaluation coefficient of the English oral intelligent interaction audio, achieving the effect of improving the accuracy of judging elision in context co-articulation changes.
[0028] The technical solution in the embodiment of the present application for solving the problem of insufficient accuracy in judging elision in context co-articulation changes is generally as follows:
[0029] First, the environmental detection and evaluation module is responsible for collecting and evaluating the interactive environment data, generating an environmental evaluation coefficient, and transmitting it to the threshold comparison module for threshold range comparison. At the same time, the audio-visual collection and processing module collects the user's oral audio data, evaluates its aspiration phenomenon, obtains an aspiration judgment evaluation coefficient, and also conducts threshold comparison. If the environmental evaluation coefficient fails to pass the threshold range comparison, the system will optimize the environment or instruct the audio-visual collection and processing module to synchronously collect lip feature data to optimize the lip features. Finally, according to the aspiration judgment evaluation coefficient of the audio, the suggestion output module provides text, audio, and visual error correction suggestions, thus forming a comprehensive, dynamic, and interactive oral English learning feedback mechanism, achieving the effect of improving the accuracy of judging aspiration in the co-articulatory changes of context.
[0030] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.
[0031] As Figure 1 shown, it is a schematic structural diagram of an intelligent interactive teaching system for oral English emotion based on deep learning provided by an embodiment of the present application. The intelligent interactive teaching system for oral English emotion based on deep learning provided by an embodiment of the present application includes: an environmental detection and evaluation module, a threshold comparison module, an audio-visual collection and processing module, an optimization module, and a suggestion output module; the environmental detection and evaluation module: used to collect and evaluate the intelligent interactive environment data of oral English, obtain an intelligent interactive environment evaluation coefficient of oral English, and transmit the intelligent interactive environment evaluation coefficient of oral English to the threshold comparison module; the threshold comparison module: used to conduct threshold range comparison and reference threshold comparison of images on the intelligent interactive environment evaluation coefficient of oral English obtained in the environmental detection and evaluation module, and conduct threshold comparison on the aspiration judgment evaluation coefficient of the intelligent interactive audio of oral English obtained in the audio-visual collection and processing module; the audio-visual collection and processing module: used to collect and evaluate the intelligent interactive audio data of oral English input by the user, obtain an aspiration judgment evaluation coefficient of the intelligent interactive audio of oral English, and transmit the aspiration judgment evaluation coefficient of the intelligent interactive audio of oral English to the suggestion output module; the optimization module: used to optimize the intelligent interactive environment of oral English according to the threshold range comparison result of the intelligent interactive environment evaluation coefficient of oral English. If the optimized intelligent interactive environment evaluation coefficient of oral English still fails to pass the threshold range comparison, send an instruction to the audio-visual collection and processing module to synchronously collect the intelligent interactive lip feature data when collecting the intelligent interactive audio data of oral English, and optimize the intelligent interactive lip features according to the threshold comparison result of the intelligent interactive environment evaluation coefficient of oral English; the suggestion output module: used to output text, audio, and visual error correction suggestions for oral English according to the aspiration judgment evaluation coefficient of the intelligent interactive audio of oral English.
[0032] Further, the threshold comparison module includes a threshold range comparison unit and a reference threshold comparison unit for the image; the threshold range comparison unit: is used to obtain the threshold range of the evaluation coefficient of the intelligent oral English interaction environment from the database, compare the evaluation coefficient of the intelligent oral English interaction environment with the threshold range of the evaluation coefficient of the intelligent oral English interaction environment. If the evaluation coefficient of the intelligent oral English interaction environment is within the threshold range of the evaluation coefficient of the intelligent oral English interaction environment, it sends an instruction to the audio-video collection and processing module to collect and evaluate the input intelligent oral English interaction audio data of the user, and obtain the explosion loss judgment evaluation coefficient of the intelligent oral English interaction audio. If the evaluation coefficient of the intelligent oral English interaction environment is greater than or equal to the maximum value of the threshold range of the evaluation coefficient of the intelligent oral English interaction environment, it sends an instruction to the optimization module to perform dynamic weight allocation for spatial positioning of the intelligent oral English interaction environment. If the evaluation coefficient of the intelligent oral English interaction environment is less than the minimum value of the threshold range of the evaluation coefficient of the intelligent oral English interaction environment, it sends an instruction to the optimization module to perform frequency domain signal processing on the intelligent oral English interaction environment. After the dynamic weight allocation for spatial positioning or the frequency domain signal processing, it sends an instruction to the environment detection and evaluation module to re-collect and evaluate the data of the intelligent oral English interaction environment, obtain the evaluation coefficient of the intelligent oral English interaction environment, and transmit the evaluation coefficient of the intelligent oral English interaction environment to the threshold comparison module. If the evaluation coefficient of the intelligent oral English interaction environment is still not within the threshold range of the evaluation coefficient of the intelligent oral English interaction environment at this time, it sends an instruction to the audio-video collection and processing module to synchronously collect the lip shape feature data of the intelligent oral English interaction when collecting the intelligent oral English interaction audio data, and perform combined analysis on the intelligent oral English interaction audio data and the optimized lip shape feature data of the intelligent oral English interaction to obtain the explosion loss judgment evaluation coefficient of the intelligent oral English interaction audio. If the evaluation coefficient of the intelligent oral English interaction environment is within the threshold range of the evaluation coefficient of the intelligent oral English interaction environment at this time, it sends an instruction to the audio-video collection and processing module to collect and evaluate the input intelligent oral English interaction audio data of the user, and obtain the explosion loss judgment evaluation coefficient of the intelligent oral English interaction audio;The reference threshold comparison unit of the said image: It is used to obtain the sharpness reference threshold of the image in the intelligent oral English interaction from the database, obtain the sharpness of the image in the intelligent oral English interaction through the sensor, compare the sharpness of the image in the intelligent oral English interaction with the reference threshold of the sharpness of the image in the intelligent oral English interaction. If the sharpness of the image in the intelligent oral English interaction is less than or equal to the reference threshold of the sharpness of the image in the intelligent oral English interaction, it sends an instruction to the optimization module to optimize the lip feature data of the intelligent oral English interaction, and combines and analyzes the intelligent oral English interaction audio data with the optimized lip feature data of the intelligent oral English interaction to obtain the explosion loss judgment evaluation coefficient of the intelligent oral English interaction audio. If the sharpness of the image in the intelligent oral English interaction is greater than the reference threshold of the sharpness of the image in the intelligent oral English interaction, it sends an instruction to the audio-video collection and processing module, combines and analyzes the intelligent oral English interaction audio data with the lip feature data of the intelligent oral English interaction to obtain the explosion loss judgment evaluation coefficient of the intelligent oral English interaction audio.;
[0033] As Figure 2 shown, it is the structural schematic diagram of the threshold comparison module in the intelligent oral English emotion intelligent interaction teaching system based on deep learning provided by the embodiment of the present application.
[0034] Further, the specific process of collecting the intelligent oral English interaction environment data is as follows: collect the intelligent oral English interaction environment data through the sensor, preprocess the intelligent oral English interaction environment data to obtain the preprocessed intelligent oral English interaction environment data, and perform preliminary denoising on the preprocessed intelligent oral English interaction environment data; the intelligent oral English interaction environment data includes the noise sound pressure level of the intelligent oral English interaction environment, the light intensity of the intelligent oral English interaction environment, and the color temperature of the intelligent oral English interaction environment.
[0035] In this embodiment, the signal-to-noise ratio measured by the sensor when the user pronounces is 20 dB, indicating that the sound quality is good. Through the fast Fourier transform (FFT), the system analyzes that the noise in the environment is mainly concentrated in the low frequency band (100 - 300 Hz), which may be generated by devices such as air conditioners or fans. The photometric sensor detects that the current light intensity is 500 lux, which is suitable for video shooting and lip recognition. At the same time, the sensor does not detect obvious stroboscopic phenomena. Normalize and convert the noise sound pressure level of the intelligent oral English interaction environment, the light intensity of the intelligent oral English interaction environment, and the color temperature of the intelligent oral English interaction environment into digital signals, and analyze the preprocessed data to obtain the intelligent oral English interaction environment evaluation coefficient of 0.95.
[0036] The noise sound pressure level of the intelligent oral English interaction environment refers to the decibel (dB) measurement of the sound pressure relative to the reference value, which can be obtained by performing spectrum analysis on the collected audio signal.
[0037] The light intensity of the intelligent oral English interaction environment refers to the brightness level of the light in the environment and can be measured using an optical sensor.
[0038] The color temperature of the intelligent oral English interaction environment is used to describe the color of the light and can be directly captured by a camera.
[0039] Furthermore, the specific process for collecting the intelligent oral English interaction audio data input by the user is as follows: collect the intelligent oral English interaction audio data through a sensor, preprocess the intelligent oral English interaction audio data to obtain the preprocessed intelligent oral English interaction audio data; the intelligent oral English interaction audio data includes the phoneme combination coverage rate of the audio and the instantaneous energy peak of the audio.
[0040] In this embodiment, the system analysis found that the user's pronunciation contained 20 different phonemes, accounting for 50% of the 40 phonemes in the English pronunciation phoneme library. When the user pronounced the word "peak", the system detected an instantaneous energy peak of 0.85, indicating a relatively high pronunciation intensity. Normalize the phoneme combination coverage rate of the audio and the instantaneous energy peak of the audio, analyze the preprocessed data, and obtain the aspiration omission judgment evaluation coefficient of the intelligent oral English interaction audio as 0.70.
[0041] The phoneme combination coverage rate of the audio refers to the proportion of the actually covered factor combinations to all possible phoneme combinations that appear within a preset time period and can be directly obtained through calculation.
[0042] The instantaneous energy peak of the audio refers to the maximum energy value of the audio signal within a preset time period and can be obtained through short-time energy analysis of the audio signal.
[0043] Furthermore, the specific method for obtaining the intelligent oral English interaction environment evaluation coefficient is as follows: obtain the weight factors of the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment from the database, arrange the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment in a time series, assign weights to the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment, compare the standard values of the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment with the real-time values of the noise sound pressure level, the light intensity, and the color temperature of the intelligent oral English interaction environment at each time series point, and obtain the intelligent oral English interaction environment evaluation coefficient after averaging processing.
[0044] In this embodiment, the specific method for obtaining the evaluation coefficient of the intelligent spoken English interaction environment is as follows:
[0045] ;
[0046] ;
[0047] In the formula, represents the evaluation coefficient of the intelligent spoken English interaction environment, which is used to evaluate the influence of the environment on the system in intelligent spoken English interaction. Set several time monitoring points, a, represents the total number of time monitoring points, represents the th time monitoring point of the noise sound pressure level of the intelligent spoken English interaction environment, represents the weight factor of the noise sound pressure level of the intelligent spoken English interaction environment, represents the noise sound pressure level of the standard intelligent spoken English interaction environment, represents the th time monitoring point of the illumination intensity of the intelligent spoken English interaction environment, represents the weight factor of the illumination intensity of the intelligent spoken English interaction environment, represents the illumination intensity of the standard intelligent spoken English interaction environment, represents the th time monitoring point of the color temperature of the intelligent spoken English interaction environment, represents the weight factor of the color temperature of the intelligent spoken English interaction environment, represents the color temperature of the standard intelligent spoken English interaction environment.
[0048] If the lighting conditions are not good, it may cause the video capture device to fail to clearly capture the lip shape features, thus indirectly affecting the synchronous analysis of the voice signal. The color temperature may cause the video image to flicker, affecting the acquisition of visual information. However, the color temperature may interfere with the user's visual experience, thus indirectly affecting the user's pronunciation.
[0049] When the system is running, obtain the mapping table of weight factors through the database, and quickly extract the corresponding weight factors according to the noise sound pressure level of the current intelligent English oral interaction environment, the illumination intensity of the intelligent English oral interaction environment, and the color temperature of the intelligent English oral interaction environment, such as the weight factor of the noise sound pressure level of the intelligent English oral interaction environment, the weight factor of the illumination intensity of the intelligent English oral interaction environment, and the weight factor of the color temperature of the intelligent English oral interaction environment. This mapping table defines a clear set of association rules, which converts the specific values of the noise sound pressure level, illumination intensity, and color temperature of the intelligent English oral interaction environment into their corresponding weight factors. Under this mechanism, whether it is to achieve a one-to-one exact match or a one-to-many relationship where multiple parameters converge into a single weight, the dynamic acquisition of weight factors can be effectively achieved.
[0050] In a specific embodiment, the data example of the evaluation coefficient of the intelligent English oral interaction environment is as follows in the table.
[0051] Table 1 Data Example Table of the Evaluation Coefficient of the Intelligent English Oral Interaction Environment
[0052]
[0053] When the weight factors of the noise sound pressure level, illumination intensity, and color temperature of the intelligent English oral interaction environment are 0.2, 0.4, and 0.4 respectively, the standard values of the noise sound pressure level, illumination intensity, and color temperature of the intelligent English oral interaction environment are 35 dB(A), 400 lx, and 6000 K respectively. Through Figure 3 and the data in Table 1, it can be seen that when the illumination intensity and color temperature of the intelligent English oral interaction environment remain unchanged, the higher the noise sound pressure level of the intelligent English oral interaction environment, the greater the evaluation coefficient of the intelligent English oral interaction environment.
[0054] Further, the specific method for obtaining the explosion loss judgment evaluation coefficient of the intelligent interactive audio in spoken English is as follows: When lip feature data needs to be combined, obtain the weight factor of the lip opening and closing distance, the weight factor of the phoneme combination coverage rate of the audio, and the weight factor of the instantaneous energy peak value of the audio from the database. Assign weights to the deviation value between the lip opening and closing distance and the reference value of the lip opening and closing distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value of the audio. Sort the lip opening and closing distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value according to the number of phonemes. Combine and process the deviation value between the lip opening and closing distance and the reference value of the lip opening and closing distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value to obtain the explosion loss judgment evaluation coefficient of the intelligent interactive audio in spoken English; When lip features do not need to be combined, obtain the weight factor of the phoneme combination coverage rate of the audio and the weight factor of the instantaneous energy peak value of the audio from the database. Assign weights to the phoneme combination coverage rate of the audio and the instantaneous energy peak value of the audio. Sort the phoneme combination coverage rate of the audio and the instantaneous energy peak value according to the number of phonemes. Combine and process the phoneme combination coverage rate of the audio and the instantaneous energy peak value to obtain the explosion loss judgment evaluation coefficient of the intelligent interactive audio in spoken English.
[0055] In this embodiment, the specific method for obtaining the explosion loss judgment evaluation coefficient of the intelligent interactive audio in spoken English is as follows:
[0056] When lip feature data does not need to be combined,
[0057] ;
[0058] ;
[0059] When lip feature data needs to be combined,
[0060] ;
[0061] ;
[0062] In the formula, represents the explosion loss judgment evaluation coefficient of the intelligent interactive audio in spoken English, which is used to evaluate the load of the software during rendering. Set several statement monitoring points. b, represents the total number of statement monitoring points. represents the lip opening and closing distance at the th statement monitoring point. represents the standard value of the lip opening and closing distance. represents the weight factor of the lip opening and closing distance. represents the phoneme combination coverage rate of the audio at the th statement monitoring point. The weight factor representing the phoneme combination coverage rate of the audio Indicates the instantaneous energy peak of the audio at the th sentence monitoring point,
[0063] If the instantaneous energy peak in the audio is too high, it may cause sound distortion, reduce the clarity and intelligibility of phonemes, thereby affecting the recognition of the phoneme combination coverage rate. In the dynamic range of the audio (i.e., the difference between the softest sound and the loudest sound), the position of the instantaneous energy peak affects the perception of the phoneme combination. If the dynamic range is too large, some phonemes may become invisible next to the strong sound.
[0064] When the system runs, obtain the mapping table of the weight factors through the database, and quickly determine the weight factors according to the current lip opening distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak of the audio, such as the weight factor of the lip opening distance, the weight factor of the phoneme combination coverage rate of the audio, and the weight factor of the instantaneous energy peak of the audio. This mapping table defines a clear set of association rules, which converts the specific values of the lip opening distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak of the audio into their corresponding weight factors. Under this mechanism, whether it is to achieve an exact one-to-one match or a one-to-many relationship where multiple parameters converge into a single weight, the dynamic acquisition of the weight factors can be effectively realized.
[0065] Furthermore, the specific optimization steps for optimizing the intelligent English oral interaction environment according to the comparison result of the threshold range of the intelligent English oral interaction environment evaluation coefficient are as follows: If the intelligent English oral interaction environment evaluation coefficient is greater than or equal to the maximum value of the threshold range of the intelligent English oral interaction environment evaluation coefficient, send an instruction to the optimization module to perform dynamic weight allocation for spatial positioning of the intelligent English oral interaction environment. If the intelligent English oral interaction environment evaluation coefficient is less than the minimum value of the threshold range of the intelligent English oral interaction environment evaluation coefficient, send an instruction to the optimization module to perform frequency domain signal processing on the intelligent English oral interaction environment.
[0066] In this embodiment, the specific steps for performing dynamic weight allocation for spatial positioning are as follows: Use a microphone array to monitor the position of the sound source in the environment. According to the position of the sound source, formulate a weight allocation strategy through the beamforming algorithm. For example, if the sound source is close to a certain microphone, the weight of this microphone can be set higher. Apply different weights to the signals received by each microphone, and synthesize the weighted signals to obtain an optimized audio signal. If the position of the sound source changes, the system will re-perform sound source positioning and weight allocation to maintain the best audio quality.
[0067] The specific steps for performing frequency-domain signal processing are as follows: collect the audio signal in the environment, perform a fast Fourier transform on the collected audio signal to convert the time-domain signal into a frequency-domain signal, analyze the frequency-domain signal to identify the frequency components of noise and user speech, apply interference algorithms such as notch filters and adaptive filters to suppress or eliminate the noise frequency components, perform an inverse Fourier transform on the processed frequency-domain signal to convert it back to a time-domain signal, and output the audio signal after frequency-domain interference processing, which has a lower noise level.
[0068] Further, the specific steps for synchronously collecting lip shape feature data for intelligent spoken English interaction audio data are as follows: obtain the lip shape data for intelligent spoken English interaction through a depth sensor, extract the lip shape feature data for intelligent spoken English interaction through a deep learning model, and use hardware triggering or software triggering to synchronize audio and video; the lip shape feature data for intelligent spoken English interaction refers to the distance between the opening and closing of the lips; hardware triggering includes using a synchronization signal generator to send a pulse signal to trigger the audio and video devices simultaneously; software triggering includes achieving timestamp alignment of the audio and video streams through the NDI protocol.
[0069] In this embodiment, the depth sensor is a device capable of measuring the distance between an object and its surface. It is usually used to capture three-dimensional space information and is used in this system to obtain the lip shape data of the user.
[0070] When the user is performing an English spoken expression, the depth sensor captures the lip movement of the user in real time, records the three-dimensional position data of the lips, preprocesses the collected lip shape data, including denoising, alignment, and normalization, uses a long short-term memory network to process the time series data, selects VGG16 as the deep learning model for feature extraction, removes the fully connected layers of the model, retains the convolutional layers and pooling layers, inputs the preprocessed lip image frames into the VGG16 model, and can locate the upper and lower boundaries of the lips by analyzing the feature maps output by the convolutional layers. Extract the pixel distance between the upper and lower boundaries as the distance between the opening and closing of the lips.
[0071] VGG16 refers to a convolutional neural network model that contains 16 weight layers, including 13 convolutional layers, 5 max pooling layers, 3 fully connected layers, and a final softmax layer for classification.
[0072] Divide the photos of different lip opening and closing distances in the historical data into a training set and a validation set at a preset ratio. Use the training set to train VGG16, and then use the validation set to validate VGG16. After training and validation, input the lip photos with different opening and closing distances collected by the audio-video collection and processing module, and output the lip opening and closing distance data.
[0073] Hardware triggering refers to using a synchronous signal generator to send pulse signals to ensure that the audio and video devices start recording simultaneously. Software triggering refers to achieving timestamp alignment of the audio-video stream through the NDI (Network Device Interface) protocol or other audio-video synchronization protocols.
[0074] Furthermore, the specific output method for text, audio, and visual English spoken language error correction suggestions based on the elision judgment evaluation coefficient of the English spoken language intelligent interaction audio:
[0075] Match the predefined text, audio, and visual English spoken language error correction suggestions from the database according to the elision judgment evaluation coefficient of the English spoken language intelligent interaction audio. The text, audio, and visual English spoken language error correction suggestions include spectrograms showing the user's pronunciation and the standard pronunciation, playing back the user's English spoken language recording and highlighting the error parts, pointing out the pronunciation error types, and providing correct written pronunciation examples.
[0076] In this embodiment, for example, when the elision judgment evaluation coefficient of the English spoken language intelligent interaction audio is less than 0.5, match the failed text, audio, and visual English spoken language error correction suggestions; when the elision judgment evaluation coefficient of the English spoken language intelligent interaction audio is greater than or equal to 0.5 and less than 1.0, match the good text, audio, and visual English spoken language error correction suggestions; when the elision judgment evaluation coefficient of the English spoken language intelligent interaction audio is greater than or equal to 1.0, match the excellent text, audio, and visual English spoken language error correction suggestions.
[0077] Further, the deep learning-based intelligent interactive teaching method for spoken English emotions is as follows: collect and evaluate the data of the intelligent interactive environment for spoken English to obtain the evaluation coefficient of the intelligent interactive environment for spoken English, and transmit the evaluation coefficient of the intelligent interactive environment for spoken English to the threshold comparison module; compare the evaluation coefficient of the intelligent interactive environment for spoken English obtained in the environment detection and evaluation module with the threshold range and the reference threshold of the image, and compare the evaluation coefficient of the burst loss judgment of the intelligent interactive audio for spoken English obtained in the audio and video collection and processing module with the threshold; collect and evaluate the intelligent interactive audio data of spoken English input by the user to obtain the evaluation coefficient of the burst loss judgment of the intelligent interactive audio for spoken English, and transmit the evaluation coefficient of the burst loss judgment of the intelligent interactive audio for spoken English to the suggestion output module; optimize the intelligent interactive environment for spoken English according to the comparison result of the threshold range of the evaluation coefficient of the intelligent interactive environment for spoken English. If the evaluation coefficient of the optimized intelligent interactive environment for spoken English still fails to pass the threshold range comparison, send an instruction to the audio and video collection and processing module to synchronously collect the lip feature data of the intelligent interactive audio for spoken English when collecting the intelligent interactive audio data, and optimize the lip feature of the intelligent interactive audio for spoken English according to the comparison result of the threshold of the evaluation coefficient of the intelligent interactive environment for spoken English; output text, audio, and visual English spoken error correction suggestions according to the evaluation coefficient of the burst loss judgment of the intelligent interactive audio for spoken English.
[0078] As Figure 4 shown, it is a schematic diagram of the deep learning-based intelligent interactive teaching method for spoken English emotions provided by the embodiment of the present application.
[0079] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0080] The present invention is described with reference to the flowcharts and / or block diagrams of systems, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for realizing in the process Figure 1 a process or multiple processes and / or blocks Figure 1a device for the functions specified in one or more boxes.
[0081] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one Figure 1 or more processes and / or boxes Figure 1 or one or more boxes.
[0082] These computer program instructions may also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 or more processes and / or boxes Figure 1 or one or more boxes.
[0083] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0084] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. The English spoken emotion intelligent interactive teaching system based on deep learning is characterized by: It includes environment detection and assessment module, threshold comparison module, audio and video collection and processing module, optimization module and suggestion output module; The environment detection and evaluation module is used to collect and evaluate the spoken English intelligent interaction environment data, obtain the spoken English intelligent interaction environment evaluation coefficient, and transmit the spoken English intelligent interaction environment evaluation coefficient to the threshold comparison module; The threshold comparison module is used to compare the English spoken intelligent interaction environment evaluation coefficient obtained in the environment detection and evaluation module with the threshold range and the reference threshold of the image, and to compare the misfire judgment evaluation coefficient of the English spoken intelligent interaction audio obtained in the audio and video collection and processing module with the threshold; The audio and video collection and processing module is used to collect and evaluate the English spoken intelligent interactive audio data input by the user, obtain the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio, and transmit the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio to the suggestion output module; The optimization module is used to optimize the spoken English intelligent interaction environment according to the comparison result of the threshold range of the spoken English intelligent interaction environment evaluation coefficient. If the optimized spoken English intelligent interaction environment evaluation coefficient still fails to pass the threshold range comparison, send a command to the audio and video collection and processing module to collect the spoken English intelligent interaction lip shape feature data synchronously when collecting the spoken English intelligent interaction audio data, and optimize the spoken English intelligent interaction lip shape feature according to the comparison result of the threshold range of the spoken English intelligent interaction environment evaluation coefficient; The suggestion output module is used to output text, audio and visual English spoken error correction suggestions according to the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio; The specific process of collecting the English spoken intelligent interactive environment data is as follows: Collecting spoken English intelligent interaction environment data through sensors, preprocessing the spoken English intelligent interaction environment data to obtain preprocessed spoken English intelligent interaction environment data, and performing preliminary noise removal on the preprocessed spoken English intelligent interaction environment data; The English speaking intelligent interaction environment data includes the noise sound pressure level of the English speaking intelligent interaction environment, the light intensity of the English speaking intelligent interaction environment, and the color temperature of the English speaking intelligent interaction environment; The specific process of collecting the English spoken intelligent interactive audio data input by the user is as follows: The English spoken intelligent interactive audio data is collected by a sensor, and the English spoken intelligent interactive audio data is preprocessed to obtain the preprocessed English spoken intelligent interactive audio data; The English spoken intelligent interactive audio data includes the phoneme combination coverage of the audio and the instantaneous energy peak of the audio.
2. The English spoken language emotional intelligent interactive teaching system based on deep learning as claimed in claim 1, characterized in that: The threshold comparison module includes a threshold range comparison unit and an image reference threshold comparison unit; The threshold range comparison unit is used to obtain the threshold range of the oral English intelligent interaction environment evaluation coefficient from the database, compare the oral English intelligent interaction environment evaluation coefficient with the threshold range of the oral English intelligent interaction environment evaluation coefficient, if the oral English intelligent interaction environment evaluation coefficient is within the threshold range of the oral English intelligent interaction environment evaluation coefficient, send a command to the audio and video collection and processing module, collect and evaluate the oral English intelligent interaction audio data input by the user, and obtain the misfire judgment evaluation coefficient of the oral English intelligent interaction audio, if the oral English intelligent interaction environment evaluation coefficient is greater than or equal to the maximum value of the threshold range of the oral English intelligent interaction environment evaluation coefficient, send a command to the optimization module, perform dynamic weight allocation for spatial positioning of the oral English intelligent interaction environment, if the oral English intelligent interaction environment evaluation coefficient is less than the minimum value of the threshold range of the oral English intelligent interaction environment evaluation coefficient, send a command to the optimization module, perform frequency domain signal processing on the oral English intelligent interaction environment, in the dynamic weight allocation or spatial positioning of the optimization module. After the frequency domain signal is processed, a command is sent to the environment detection and evaluation module to collect and evaluate the spoken English intelligent interaction environment data again to obtain the spoken English intelligent interaction environment evaluation coefficient, and the spoken English intelligent interaction environment evaluation coefficient is transmitted to the threshold comparison module. If the spoken English intelligent interaction environment evaluation coefficient is still not within the threshold range of the spoken English intelligent interaction environment evaluation coefficient at this time, a command is sent to the audio and video collection and processing module to synchronously collect the spoken English intelligent interaction lip shape feature data when collecting the spoken English intelligent interaction audio data, and the spoken English intelligent interaction audio data is combined with the optimized spoken English intelligent interaction lip shape feature data for analysis to obtain the misfire judgment evaluation coefficient of the spoken English intelligent interaction audio. If the spoken English intelligent interaction environment evaluation coefficient is within the threshold range of the spoken English intelligent interaction environment evaluation coefficient at this time, a command is sent to the audio and video collection and processing module to collect and evaluate the spoken English intelligent interaction audio data input by the user to obtain the misfire judgment evaluation coefficient of the spoken English intelligent interaction audio. The reference threshold comparison unit of the image is used to obtain the sharpness reference threshold of the image in the intelligent interaction of spoken English from the database, obtain the sharpness of the image in the intelligent interaction of spoken English through the sensor, compare the sharpness of the image in the intelligent interaction of spoken English with the reference threshold of the sharpness of the image in the intelligent interaction of spoken English, if the sharpness of the image in the intelligent interaction of spoken English is less than or equal to the reference threshold of the sharpness of the image in the intelligent interaction of spoken English, send an instruction to the optimization module, optimize the lip shape feature data of the intelligent interaction of spoken English, combine and analyze the audio data of the intelligent interaction of spoken English with the optimized lip shape feature data of the intelligent interaction of spoken English, and obtain the misfire judgment evaluation coefficient of the audio of the intelligent interaction of spoken English; if the sharpness of the image in the intelligent interaction of spoken English is greater than the reference threshold of the sharpness of the image in the intelligent interaction of spoken English, send an instruction to the audio and video collection and processing module, combine and analyze the audio data of the intelligent interaction of spoken English with the lip shape feature data of the intelligent interaction of spoken English, and obtain the misfire judgment evaluation coefficient of the audio of the intelligent interaction of spoken English.
3. The English spoken language emotional intelligence interactive teaching system based on deep learning as claimed in claim 1, characterized in that: The specific method for obtaining the evaluation coefficient of the English oral intelligent interactive environment is: The weight factors of the noise sound pressure level, the light intensity and the color temperature of the spoken English intelligent interaction environment are obtained from the database, the noise sound pressure level, the light intensity and the color temperature of the spoken English intelligent interaction environment are arranged in time series, weights are assigned to the noise sound pressure level, the light intensity and the color temperature of the spoken English intelligent interaction environment, the standard values of the noise sound pressure level, the light intensity and the color temperature of the spoken English intelligent interaction environment at each time series point are compared with the real-time values of the noise sound pressure level, the light intensity and the color temperature of the spoken English intelligent interaction environment, and the evaluation coefficient of the spoken English intelligent interaction environment is obtained after averaging.
4. The English spoken language emotional intelligence interactive teaching system based on deep learning as claimed in claim 1, characterized in that: The specific method for obtaining the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio is: If it is necessary to combine the lip feature data, obtain the weight factor of the lip opening and closing distance, the weight factor of the phoneme combination coverage rate of the audio, and the weight factor of the instantaneous energy peak of the audio from the database, assign weights to the deviation value of the lip opening and closing distance from the lip opening and closing distance reference value, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value of the audio, sort the lip opening and closing distance, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value of the audio according to the number of phonemes, combine and process the deviation value of the lip opening and closing distance from the lip opening and closing distance reference value, the phoneme combination coverage rate of the audio, and the instantaneous energy peak value of the audio, and obtain the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio; If it is not necessary to combine lip features, obtain the weight factor of the phoneme combination coverage rate of the audio and the weight factor of the instantaneous energy peak of the audio from the database, assign weights to the phoneme combination coverage rate of the audio and the instantaneous energy peak of the audio, sort the phoneme combination coverage rate of the audio and the instantaneous energy peak of the audio according to the number of phonemes, combine and process the phoneme combination coverage rate of the audio and the instantaneous energy peak of the audio to obtain the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio.
5. The English spoken language emotional intelligence interactive teaching system based on deep learning as claimed in claim 1, characterized in that: The specific optimization steps for optimizing the spoken English intelligent interaction environment according to the comparison result of the threshold range of the spoken English intelligent interaction environment evaluation coefficient are as follows: If the evaluation coefficient of the spoken English intelligent interaction environment is greater than or equal to the maximum value of the threshold range of the spoken English intelligent interaction environment evaluation coefficient, an instruction is sent to the optimization module to perform dynamic weight allocation for spatial positioning of the spoken English intelligent interaction environment; if the evaluation coefficient of the spoken English intelligent interaction environment is less than the minimum value of the threshold range of the spoken English intelligent interaction environment evaluation coefficient, an instruction is sent to the optimization module to perform frequency domain signal processing on the spoken English intelligent interaction environment.
6. The English spoken language emotional intelligence interactive teaching system based on deep learning as claimed in claim 1, characterized in that: The specific steps of synchronously collecting the English spoken intelligent interactive lip shape feature data when collecting the English spoken intelligent interactive audio data are as follows: Acquire English spoken intelligent interactive lip shape data through depth sensors, extract English spoken intelligent interactive lip shape feature data through deep learning models, and use hardware triggers or software triggers to synchronize audio and video; The English spoken intelligent interactive lip shape feature data includes the lip opening and closing distance; Hardware triggering involves using a sync signal generator to send a pulse signal to trigger both audio and video devices simultaneously; Software triggering includes aligning audio and video stream timestamps through the NDI protocol.
7. The deep learning-based English oral emotional intelligent interactive teaching system according to claim 1, characterized in that: The specific output method of outputting text, audio and visual English spoken error correction suggestions according to the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio: According to the misfire judgment evaluation coefficient of the intelligent interactive audio of spoken English, predefined text, audio and visual spoken English error correction suggestions are matched from the database. The text, audio and visual spoken English error correction suggestions include displaying the spectrogram of the user's pronunciation and the standard pronunciation, playing back the user's spoken English recording and highlighting the error parts, pointing out the type of pronunciation error and providing correct text pronunciation examples.
8. The method applied to the deep learning-based English spoken emotional intelligence interactive teaching system according to any one of claims 1 to 7, characterized in that: Collect and evaluate the spoken English intelligent interaction environment data, obtain the spoken English intelligent interaction environment evaluation coefficient, and transmit the spoken English intelligent interaction environment evaluation coefficient to the threshold comparison module; The English spoken intelligent interactive environment evaluation coefficient obtained in the environment detection and evaluation module is compared with the threshold range and the reference threshold of the image, and the English spoken intelligent interactive audio misfire judgment evaluation coefficient obtained in the audio and video collection and processing module is compared with the threshold; Collect and evaluate the English spoken intelligent interactive audio data input by the user, obtain the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio, and transmit the misfire judgment evaluation coefficient of the English spoken intelligent interactive audio to the suggestion output module; The English speaking intelligent interaction environment is optimized according to the comparison result of the threshold range of the evaluation coefficient of the English speaking intelligent interaction environment. If the optimized evaluation coefficient of the English speaking intelligent interaction environment still fails to pass the comparison of the threshold range, a command is sent to the audio and video collection and processing module to collect the English speaking intelligent interaction lip shape feature data synchronously when collecting the English speaking intelligent interaction audio data, and the English speaking intelligent interaction lip shape feature is optimized according to the comparison result of the threshold range of the English speaking intelligent interaction environment evaluation coefficient; According to the misfire judgment evaluation coefficient of the intelligent interactive audio of spoken English, text, audio and visual English speaking error correction suggestions are output.
Citation Information
Patent Citations
A method to improve pronunciation quality in English teaching
CN111916106B
A method, apparatus, device, and storage medium for evaluating plosive sounds.
CN113077822B
Interactive virtual teacher system having intelligent error correction function
CN102169642A
Howling detection method and device, storage medium and electronic equipment
CN113271386A
Intelligent speech recognition and sentiment analysis system and method
CN118197303A