AI-based English listening and speaking ability comprehensive improvement method and system

Through deep learning models and gamified English teaching system, the problems of inaccurate pronunciation recognition and insufficient interaction in the existing technology are solved, and personalized real-time feedback and interactive exercises are achieved for non-native speakers, which significantly improves the learning effect.

CN120564490APending Publication Date: 2025-08-29GUANGXI INST OF ELECTROMECHANICAL TECHNICIANS (GUANGXI SENIOR MECHANICAL TECH SCHOOL GUANGXI ADULT ELECTROMECHANICAL SECONDARY VOCATIONAL SCHOOL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510640873.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing English teaching system lacks accuracy in pronunciation recognition of non-native speakers, lacks real-time feedback and interactivity, resulting in low user participation and poor learning results.

Method used

A speech recognition model combined with deep neural network and convolutional neural network is adopted to recognize and feedback pronunciation errors in real time, and combine personalized learning paths and gamified design to provide highly interactive oral practice scenarios.

Benefits of technology

It significantly improves the accuracy of pronunciation recognition, provides personalized learning paths and highly interactive exercises, and improves users' learning interest and effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564490A_ABST
    Figure CN120564490A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of English teaching, and discloses an AI-based English listening and speaking ability comprehensive improvement method and system, and the method comprises the steps: carrying out the voice recognition and real-time feedback, carrying out the real-time recognition and analysis of the spoken language of a user through an advanced AI voice recognition technology, and carrying out the deep learning of the spoken language of the user through a deep learning model. Especially, optimization is carried out aiming at the pronunciation characteristics of a non-native language user; the learning content and the practice difficulty are dynamically adjusted according to the learning progress and the weak links of the user by using an AI algorithm, and a personalized learning path is provided; a spoken language practice scene with high interactivity is designed, and gamification elements are introduced. According to the invention, more accurate speech recognition and real-time feedback, a personalized learning path, high-interactivity practice and multi-modal learning resources can be provided, so that the English listening and speaking ability of the user is comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to, but is not limited to, the field of English teaching technology, and in particular relates to an AI-based method and system for comprehensively improving English listening and speaking skills. Background Art

[0002] English teaching is responsible for cultivating students' basic English literacy and developing their thinking ability. That is, students master basic English language knowledge through English courses, develop basic English listening, speaking, reading and writing skills, and initially form the ability to communicate with others in English, further promote the development of thinking ability, and lay the foundation for continuing to learn English and learning other related scientific and cultural knowledge in English in the future. In the process of English teaching, oral interaction is often required through practicing English dialogue scenarios.

[0003] Existing technology uses recording and text conversion technology, allowing students to record themselves speaking English. The system then converts the recording into text, compares it with standard answers, and provides feedback. Representative systems of this approach include: Google Voice Typing, which uses speech recognition technology to convert users' spoken words into text; and Speechace, an API service that provides oral assessment, capable of recognizing speech and scoring pronunciation. This method suffers from limited speech recognition accuracy, particularly for non-native speakers' accents and pronunciation errors; it lacks real-time feedback and correction suggestions; and it lacks interactivity, resulting in low user engagement.

[0004] Another common technique is video-based instruction, which aims to improve English listening and speaking skills by watching English instructional videos and practicing with them. Representative systems include Rosetta Stone, which uses a combination of video, images, and sound for language instruction; and Duolingo, which provides rich exercises and interactive video content to improve language skills. This approach offers fixed video instruction content, lacking personalized and targeted training. User engagement and interactivity are limited, with a lack of real-time practice and feedback. Furthermore, the videos offer limited feedback on oral practice, relying primarily on user self-assessment.

[0005] In view of the above analysis, the technical problems that need to be solved urgently in the existing technology are:

[0006] Existing technologies have limited accuracy, limited user interactivity, and lack of real-time practice and feedback. Summary of the Invention

[0007] In response to the problems existing in the prior art, the present invention provides an AI-based method and system for comprehensively improving English listening and speaking skills.

[0008] The present invention is implemented as follows: a method for comprehensively improving English listening and speaking skills based on AI, the method specifically comprising:

[0009] S1: Speech recognition and real-time feedback. This system uses advanced AI speech recognition technology to identify and analyze users' spoken language in real time. Through deep learning models, it optimizes the pronunciation characteristics of non-native speakers, including but not limited to phoneme recognition, pitch analysis, and speech rate detection. It also provides instant corrections and suggestions when users are speaking.

[0010] S2: Leveraging AI algorithms, we dynamically adjust learning content and exercise difficulty based on the user's learning progress and weaknesses, providing a personalized learning path. Specifically, we analyze data to identify common mistakes and areas that need improvement, generate personalized practice tasks and learning materials, and continuously update the learning path based on user performance.

[0011] S3: Design highly interactive oral practice scenarios and introduce gamification elements, including setting up virtual dialogue scenarios, role-playing and challenging tasks, stimulating users' learning interest and enthusiasm through points, rewards and leaderboards, and providing a variety of interactive practice forms, such as simulated daily conversations, situational dialogues and oral practice on specific topics, to enhance users' practical application capabilities.

[0012] Furthermore, the S1 specifically includes:

[0013] (1) Data collection and preprocessing: Collect a large amount of English spoken data, covering the pronunciation of non-native speakers from different countries and regions, clean the data, remove noise and irrelevant parts, enhance the data quality, and annotate the data, including the types of pronunciation errors and accent characteristics;

[0014] (2) Model training: A speech recognition model that combines a deep neural network (DNN) and a convolutional neural network (CNN) is used for training. Labeled data is used for training. The model is able to recognize different accents and pronunciation errors. The training data is input into the DNN for feature extraction, and then processed by the CNN for high-level feature processing. Using transfer learning methods, the existing speech recognition model is further fine-tuned to adapt to specific user groups.

[0015] (3) Real-time recognition and feedback: Users record spoken language through the system, which recognizes and transcribes the speech in real time, compares the recognition results with the standard pronunciation, analyzes the location and type of pronunciation errors, and provides real-time feedback to users on the specific location of the pronunciation errors and improvement suggestions.

[0016] Furthermore, the S2 specifically includes:

[0017] (1) Data collection and analysis: Collect users’ historical learning data, including learning progress, practice results, error types, etc., analyze users’ learning patterns, and identify weak links and areas for improvement;

[0018] (2) Personalized learning plan generation: Use recommendation algorithms (such as collaborative filtering, content recommendation, etc.) to generate personalized learning plans based on the user's learning data and dynamically adjust the learning content and difficulty;

[0019] (3) Implementation and feedback: Based on the personalized learning plan, recommend appropriate learning content and exercises to users, regularly evaluate users' learning effects, and dynamically adjust the learning path.

[0020] Furthermore, step S2 includes the following specific steps:

[0021] S21: By collecting users' historical learning data and real-time practice performance, we build user learning profiles, which include data on multiple dimensions such as pronunciation accuracy, grammatical error rate, vocabulary, and oral fluency;

[0022] S22: Utilize AI algorithms to analyze user learning profiles and identify weaknesses at different learning stages, with a particular focus on common errors in pronunciation, grammar, and vocabulary.

[0023] S23: Based on the analysis results, dynamically generate personalized learning plans and adjust the difficulty and type of learning content, including adding specialized exercises targeting users' weak points, such as pronunciation correction, grammar practice, and vocabulary memorization;

[0024] S24: Through dynamic adjustments, users’ learning profiles are regularly updated to ensure that the learning content matches the user’s current level and learning progress. The learning path is optimized and adjusted in real time based on the user’s learning performance and feedback.

[0025] S25: Provides a variety of feedback mechanisms, including detailed error analysis reports, targeted improvement suggestions and progress tracking charts, so that users can clearly understand their learning progress and areas for improvement, thereby improving learning effectiveness and efficiency.

[0026] The personalized learning path designed in this way can more effectively meet the user's learning needs and improve the user's learning interest and effect.

[0027] Furthermore, the S3 specifically includes:

[0028] (1) Designing interactive practice scenarios: Create a variety of oral practice scenarios, such as role-playing and dialogue simulation, and use reinforcement learning algorithms to enhance interactivity;

[0029] (2) Introduction of gamification elements: Design gamification elements such as points, levels, and rewards to motivate users to continue learning. Set different tasks and challenges, and users will receive corresponding rewards and points after completing the tasks.

[0030] (3) Implementation and optimization: Implement interactive and gamified exercise content, collect user feedback and data, and continuously optimize the exercise content and gamification design based on user feedback and data.

[0031] Another object of the present invention is to provide an AI-based comprehensive English listening and speaking ability improvement system, which specifically includes:

[0032] The speech recognition and real-time feedback module is used to recognize and analyze the user's spoken language in real time, immediately provide feedback on the specific location of pronunciation errors, and provide correct pronunciation demonstrations and improvement suggestions;

[0033] Personalized learning path module, which dynamically adjusts learning content and exercise difficulty based on the user's learning progress and weaknesses, providing a personalized learning path;

[0034] The interactivity and gamification design module is used to design highly interactive oral practice scenarios and introduce gamification elements.

[0035] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0036] First, this invention significantly improves the recognition accuracy of different accents and pronunciation errors, especially for non-native speakers, by using a speech recognition model that combines a deep neural network (DNN) and a convolutional neural network (CNN). A real-time feedback mechanism quickly identifies pronunciation errors and provides improvement suggestions, improving the user's spoken English.

[0037] This invention uses collaborative filtering algorithms to generate personalized learning plans based on the user's learning progress and weaknesses, and dynamically adjusts the learning content and difficulty. This provides personalized and targeted training, improves learning efficiency, and avoids the boredom caused by fixed content.

[0038] This invention introduces highly interactive practice scenarios, such as role-playing and dialogue simulation, to increase the realism and fun of the practice. Through gamification design, it motivates users to continue learning and maintain high participation.

[0039] Second, this invention addresses the issues of insufficient personalization, delayed real-time feedback, and a monotonous learning process in existing methods for improving English listening and speaking skills. Traditional English learning systems often use fixed learning paths and content, making it difficult to dynamically adjust to each user's specific learning needs, resulting in unsatisfactory learning results. Furthermore, existing speech recognition technology is insufficiently optimized for the pronunciation characteristics of non-native speakers, resulting in inaccurate feedback and a negative impact on the user's learning experience.

[0040] This invention utilizes advanced AI speech recognition technology and deep learning models to optimize the pronunciation characteristics of non-native speakers, achieving real-time recognition and feedback of users' spoken language. Users receive immediate and accurate corrections and suggestions during pronunciation, significantly improving the effectiveness and efficiency of pronunciation practice. Furthermore, by dynamically adjusting learning content and difficulty, it provides personalized learning paths, ensuring targeted training and improvement at every stage of learning, significantly improving learning efficiency.

[0041] Furthermore, the system incorporates highly interactive oral practice scenarios and gamification elements, stimulating user interest and motivation through diverse interactive forms such as virtual conversations, role-playing, and challenging tasks. Gamification mechanisms such as points, rewards, and leaderboards further enhance the fun of learning, allowing users to improve their English listening and speaking skills in a relaxed and enjoyable environment, overcoming the monotony of traditional learning methods.

[0042] In summary, this invention achieves significant technological advancements while resolving existing technical issues. By combining AI technology, deep learning models, and gamification, this invention not only enhances learning personalization and real-time feedback, but also makes the learning process more interactive and engaging. This provides users with a brand-new solution for improving their English listening and speaking skills, significantly improving learning outcomes and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a flow chart of a method for comprehensively improving English listening and speaking skills based on AI provided by an embodiment of the present invention;

[0044] Figure 2 This is a module diagram of an AI-based comprehensive English listening and speaking ability improvement system provided by an embodiment of the present invention;

[0045] Figure 3 is a speech recognition accuracy curve provided by an embodiment of the present invention;

[0046] Figure 4 This is a user learning effect improvement curve provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0048] like Figure 1 As shown, an embodiment of the present invention provides an AI-based method for comprehensively improving English listening and speaking skills, which specifically includes:

[0049] S1: Speech recognition and real-time feedback, using advanced AI speech recognition technology to identify and analyze users' spoken language in real time. Through deep learning models, it is optimized specifically for the pronunciation characteristics of non-native speakers.

[0050] S2: Leveraging AI algorithms, we dynamically adjust learning content and exercise difficulty based on the user's learning progress and weaknesses, providing a personalized learning path.

[0051] S3: Design interactive oral practice scenarios and introduce gamification elements.

[0052] The S1 uses a speech recognition model that combines a deep neural network (DNN) and a convolutional neural network (CNN) to improve the recognition accuracy of different accents and pronunciation errors. It analyzes the user's pronunciation in real time, compares it with the standard pronunciation, and provides instant feedback on the specific location of the pronunciation error and improvement suggestions, including:

[0053] (4) Data collection and preprocessing

[0054] Collect a large amount of English spoken data, covering the pronunciation of non-native speakers from different countries and regions, clean the data, remove noise and irrelevant parts, enhance the data quality, and annotate the data, including the types of pronunciation errors and accent characteristics.

[0055] (2) Model training

[0056] This speech recognition model uses a combination of a deep neural network (DNN) and a convolutional neural network (CNN). Trained using labeled data, the model can identify different accents and mispronunciations. The training data is fed into the DNN for feature extraction, and then processed by the CNN for high-level feature processing.

[0057] Specific model architecture:

[0058] Input layer: Input speech signal and perform preprocessing (such as MFCC feature extraction).

[0059] DNN layer: multi-layer fully connected layer for feature extraction.

[0060] CNN layer: multiple layers of convolutional layers and pooling layers for high-level feature extraction.

[0061] Output layer: Softmax layer, used for classification output.

[0062]

[0063] Use transfer learning methods to further fine-tune the existing speech recognition model to adapt it to specific user groups.

[0064] (3) Real-time recognition and feedback

[0065] Users record spoken language through the system, which then recognizes and transcribes it in real time. The system compares the recognition results with standard pronunciation to analyze the location and type of pronunciation errors. The system then provides real-time feedback to users on the specific location of pronunciation errors and suggestions for improvement.

[0066]

[0067] Said S2 specifically includes:

[0068] (1) Data collection and analysis

[0069] Collect users' historical learning data, including learning progress, practice results, error types, etc., analyze users' learning patterns, and identify weak links and areas for improvement.

[0070]

[0071]

[0072] (2) Personalized learning plan generation

[0073] Use recommendation algorithms (such as collaborative filtering, content recommendation, etc.) to generate personalized learning plans based on users' learning data, dynamically adjust learning content and difficulty, and ensure that users continue to make progress with appropriate challenges.

[0074] (3) Implementation and feedback

[0075] Based on personalized learning plans, we recommend suitable learning content and exercises to users, regularly evaluate users’ learning outcomes, and dynamically adjust learning paths to ensure users’ continuous progress.

[0076] Said S3 specifically includes:

[0077] (1) Design interactive practice scenarios

[0078] Create a variety of oral practice scenarios, such as role-playing and dialogue simulations, use reinforcement learning algorithms to enhance interactivity, increase the realism and fun of the practice, and ensure that the practice content is rich and diverse, covering different scenarios such as daily conversations and business exchanges.

[0079]

[0080]

[0081] (2) Introduction of gamification elements

[0082] Design gamification elements such as points, levels, and rewards to motivate users to continue learning. Set different tasks and challenges, and users will receive corresponding rewards and points after completing the tasks.

[0083]

[0084]

[0085] (3) Implementation and optimization

[0086] Implement interactive and gamified exercises and collect user feedback and data. Based on this feedback and data, continuously optimize the exercise content and gamification design to enhance user experience and learning outcomes.

[0087] like Figure 2 As shown, an embodiment of the present invention provides an AI-based comprehensive English listening and speaking ability improvement system, which specifically includes:

[0088] The speech recognition and real-time feedback module is used to recognize and analyze the user's spoken language in real time, immediately provide feedback on the specific location of pronunciation errors, and provide correct pronunciation demonstrations and improvement suggestions;

[0089] Personalized learning path module, which dynamically adjusts learning content and exercise difficulty based on the user's learning progress and weaknesses, providing a personalized learning path;

[0090] The interactivity and gamification design module is used to design highly interactive oral practice scenarios and introduce gamification elements.

[0091] 1. Specific application fields or related products of the present invention.

[0092] Example 1: Interactive Speech Correction and Pronunciation Challenge Application

[0093] Application Background:

[0094] To help non-native English learners improve their spoken pronunciation, we developed an interactive speech correction and pronunciation challenge app. This app combines advanced speech recognition technology with personalized learning paths, while also incorporating gamification to increase learning interest and motivation.

[0095] Specific implementation:

[0096] 1. Voice recognition and real-time feedback module:

[0097] Users record their own pronunciation through the app, and the app's built-in DNNCNN hybrid speech recognition model analyzes the user's pronunciation in real time and compares it with the pronunciation in the standard pronunciation library.

[0098] The system immediately provides feedback on the specific location of pronunciation errors (such as syllables and stress placement), and provides correct pronunciation demonstrations and improvement suggestions. Users can repeat the practice based on the feedback until their pronunciation reaches the standard.

[0099] 2. Personalized learning path module:

[0100] The app assesses the user's pronunciation level and weaknesses based on their first test score and subsequent practice performance.

[0101] Based on the evaluation results, the app recommends personalized pronunciation practice courses, from basic phonetic symbols to advanced linked reading, weak reading and other techniques, to gradually improve the user's pronunciation ability.

[0102] The learning path is dynamically adjusted to increase difficulty or adjust exercise content based on the user's progress, ensuring that users always learn in a challenge that suits them.

[0103] 3. Interactivity and gamification design:

[0104] The app features multiple pronunciation challenge levels, each containing pronunciation exercises for words, phrases, or sentences on different themes.

[0105] Users will receive points and rewards after completing each level. When the points accumulate to a certain amount, they can unlock new levels or obtain virtual badges.

[0106] The introduction of a leaderboard feature allows users to view their rankings globally or within specific user groups, stimulating a sense of competition and motivation for continuous learning.

[0107] Example 2: Role-playing oral learning platform

[0108] Application Background:

[0109] To improve users' actual oral communication skills, we developed a role-playing oral learning platform. This platform simulates real-life conversation scenarios, combines AI-assisted role-playing with real-time feedback, and helps users practice speaking in a simulated environment.

[0110] Specific implementation:

[0111] 1. AI role-playing module:

[0112] The platform provides multiple virtual characters, each with different personalities, backgrounds, and conversation styles. Users can choose to practice conversations with a specific character.

[0113] AI technology enables virtual characters to intelligently respond to user input, generate natural and fluent conversation content, and simulate real communication scenarios.

[0114] 2. Voice recognition and real-time feedback:

[0115] During the conversation, the platform recognizes the user's pronunciation and grammar in real time, providing instant feedback and suggestions. For pronunciation errors, the system will point out the specific location and provide a demonstration of correct pronunciation; for grammatical errors, the system will provide the correct sentence structure and usage explanation.

[0116] Users can adjust their expressions based on feedback and practice repeatedly until they achieve satisfactory results.

[0117] 3. Personalized learning path and progress tracking:

[0118] The platform dynamically generates personalized learning plans based on users' conversational performance and learning goals, gradually improving users' oral communication skills from daily conversations to discussions in professional fields.

[0119] It provides a learning progress tracking function, and users can view their learning reports at any time, including changes in indicators such as pronunciation improvement and conversation fluency.

[0120] 4. Interactivity and social functions:

[0121] Users can participate in role-playing exercises with learners around the world and improve their speaking skills through online collaboration and competition.

[0122] The platform has forums and communities where users can share learning experiences, exchange experiences, and even organize voice communication activities to enhance the interactivity and fun of learning.

[0123] 5. Multimodal learning resources

[0124] Provides a wealth of multimodal learning resources, including text, audio, video, and interactive exercises, to meet diverse learning needs. By integrating various learning resources, users can choose different learning methods based on their preferences and needs. Interactive video exercises allow users to practice speaking while watching videos and receive real-time feedback.

[0125] 2. Relevant evidence of the technical effects obtained by the embodiments of the present invention.

[0126] like Figure 3 As shown in the figure, the horizontal axis is the number of iterations and the vertical axis is the speech recognition accuracy (percentage). The curve shows how the speech recognition accuracy changes with the number of iterations, reflecting the performance improvement of the deep learning model during the training process. Figure 3 It demonstrates the continuous improvement of the accuracy of the speech recognition model during iterative training, proving the effectiveness of deep learning models in processing non-native pronunciations.

[0127] like Figure 4As shown, the horizontal axis is the learning time (hours) and the vertical axis is the learning effect score (score). The curve chart shows the improvement of the user's learning effect after using personalized learning paths and interactive exercises. The score is based on the user's oral proficiency test results. Figure 4 It shows that as users' learning time increases, their learning effect scores gradually improve, verifying the effectiveness of personalized learning paths and interactive exercises in improving users' English listening and speaking skills.

[0128] Through these two curve charts of technical effects, we can clearly see that the improved solution has significantly improved speech recognition accuracy and user learning effects.

[0129] These charts visually demonstrate the effectiveness of the technology solution, providing users with a more efficient, personalized and interactive learning experience.

[0130] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVDROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on a carrier medium such as a disk, CD or DVDROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., or can be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0131] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A comprehensive method for improving English listening and speaking skills based on AI, characterized by: The method specifically includes: S1: Speech recognition and real-time feedback. This system uses advanced AI speech recognition technology to identify and analyze users' spoken language in real time. Through deep learning models, it optimizes the pronunciation characteristics of non-native speakers, including but not limited to phoneme recognition, pitch analysis, and speech rate detection. It also provides instant corrections and suggestions when users are speaking. S2: Leveraging AI algorithms, we dynamically adjust learning content and exercise difficulty based on the user's learning progress and weaknesses, providing a personalized learning path. Specifically, we analyze data to identify common mistakes and areas that need improvement, generate personalized practice tasks and learning materials, and continuously update the learning path based on user performance. S3: Design highly interactive oral practice scenarios and introduce gamification elements, including setting up virtual dialogue scenarios, role-playing and challenging tasks, stimulating users' learning interest and enthusiasm through points, rewards and leaderboards, and providing a variety of interactive practice forms, including simulated daily conversations, situational dialogues and oral practice on specific topics, to enhance users' practical application capabilities.

2. The AI-based comprehensive improvement method for English listening and speaking skills according to claim 1, characterized in that: Said S1 specifically includes: (1) Data collection and preprocessing: Collect a large amount of English spoken data, covering the pronunciation of non-native speakers from different countries and regions, clean the data, remove noise and irrelevant parts, enhance the data quality, and annotate the data, including the types of pronunciation errors and accent characteristics; (2) Model training: A speech recognition model that combines a deep neural network (DNN) and a convolutional neural network (CNN) is used for training. The model is able to recognize different accents and pronunciation errors. The training data is input into the DNN for feature extraction, and then processed by the CNN for high-level feature processing. Using the transfer learning method, the existing speech recognition model is further fine-tuned to adapt to specific user groups. (3) Real-time recognition and feedback: Users record spoken language through the system, which recognizes and transcribes the speech in real time, compares the recognition results with the standard pronunciation, analyzes the location and type of pronunciation errors, and provides real-time feedback to users on the specific location of the pronunciation errors and improvement suggestions.

3. The AI-based comprehensive improvement method for English listening and speaking skills according to claim 1, characterized in that: The step S2 further includes the following specific steps: S21: By collecting users' historical learning data and real-time practice performance, we build user learning profiles, including data on multiple dimensions such as pronunciation accuracy, grammatical error rate, vocabulary, and oral fluency. S22: Use AI algorithms to analyze user learning profiles and identify weaknesses at different learning stages, with a particular focus on common errors in pronunciation, grammar, and vocabulary. S23: Based on the analysis results, dynamically generate personalized learning plans and adjust the difficulty and type of learning content, including adding specialized exercises targeting users' weak points; S24: Through dynamic adjustments, users’ learning profiles are regularly updated to ensure that the learning content matches the user’s current level and learning progress. The learning path is optimized and adjusted in real time based on the user’s learning performance and feedback. S25: Provides a variety of feedback mechanisms, including detailed error analysis reports, targeted improvement suggestions and progress tracking charts, so that users can clearly understand their learning progress and areas for improvement, thereby improving learning effectiveness and efficiency.

4. The AI-based comprehensive improvement method for English listening and speaking skills according to claim 3, characterized in that: S21: By collecting the user's historical learning data and real-time practice performance, we build a user learning profile, which includes data on multiple dimensions, such as pronunciation accuracy, grammatical error rate, vocabulary, and oral fluency. We use a vector to represent the user's learning status, where the vector elements represent the learning indicators of each dimension, which are expressed as: S22: Use AI algorithms to analyze user learning profiles and identify users’ weaknesses at different learning stages. Construct a loss function \(L\) for user learning progress. The loss function can be defined as: S23: Based on the analysis results, dynamically generate a personalized learning plan and adjust the difficulty and type of learning content. Minimize the loss function through gradient descent and update the parameters of the learning plan: S24: Through dynamic adjustment, the user's learning profile is regularly updated to ensure that the learning content matches the user's current level and learning progress. The learning path is optimized and adjusted in real time based on the user's learning performance and feedback. A recurrent neural network (RNN) is used to predict the user's learning status. The formula is as follows: S25: Provide a variety of feedback mechanisms, including detailed error analysis reports, targeted improvement suggestions, and progress tracking charts. Use regression analysis models to generate user progress charts and improvement suggestions. The formula is as follows: Through the above methods, the learning content and path can be dynamically optimized to comprehensively improve users' English listening and speaking abilities.

5. The AI-based comprehensive improvement method for English listening and speaking skills according to claim 1, characterized in that: Said S2 specifically includes: (1) Data collection and analysis: Collect users’ historical learning data, including learning progress, practice results, and error types, analyze users’ learning patterns, and identify weak links and areas for improvement; (2) Personalized learning plan generation: Using recommendation algorithms, a personalized learning plan is generated based on the user's learning data, and the learning content and difficulty are dynamically adjusted; (3) Implementation and feedback: Based on the personalized learning plan, recommend appropriate learning content and exercises to users, regularly evaluate users' learning effects, and dynamically adjust the learning path.

6. The AI-based comprehensive improvement method for English listening and speaking skills according to claim 1, characterized in that: Said S3 specifically includes: (1) Designing highly interactive practice scenarios: Create a variety of oral practice scenarios, such as role-playing and dialogue simulations, and use reinforcement learning algorithms to enhance interactivity; (2) Introduction of gamification elements: Designing gamification elements such as points, levels, and rewards to motivate users to continue learning. Setting different tasks and challenges, and users will receive corresponding rewards and points after completing tasks. (3) Implementation and optimization: Implement interactive and gamified exercise content, collect user feedback and data, and continuously optimize the exercise content and gamification design based on user feedback and data.

7. The AI-based comprehensive improvement method for English listening and speaking skills according to claim 1, characterized in that: The system specifically includes: The speech recognition and real-time feedback module is used to recognize and analyze the user's spoken language in real time, immediately provide feedback on the specific location of pronunciation errors, and provide correct pronunciation demonstrations and improvement suggestions; Personalized learning path module, which dynamically adjusts learning content and exercise difficulty based on the user's learning progress and weaknesses, providing a personalized learning path; The interactivity and gamification design module is used to design highly interactive oral practice scenarios and introduce gamification elements.