Intelligent voice companion education system and method

By adjusting voice parameters and generating knowledge lists, the problems of inaccurate playback and unclear learning decisions in existing technologies have been solved, realizing personalized and dynamically optimized intelligent voice-guided education, and improving user satisfaction and learning outcomes.

CN120636398BActive Publication Date: 2026-04-07NANJING XIAOZHUANG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies do not compensate for playback errors due to recognition vulnerabilities such as user-generated dialects, which can easily lead to incorrect playback. They also do not compensate for deviations in playback progress based on actual playback conditions, which is detrimental to playback accuracy. Furthermore, they do not make decisions on learning frequency based on educational outcomes, which presents certain limitations.

Method used

By adjusting voice parameters according to user instructions, obtaining user feedback information to calculate speech recognition accuracy and playback inclusiveness, updating the speech recognition model and preference information, generating knowledge lists and exercises, dynamically adjusting learning content, and providing a personalized and dynamically optimized educational experience.

Benefits of technology

It improves the accuracy and comfort of voice playback, optimizes the learning experience, enhances the flexibility and convenience of learning, improves learning efficiency, avoids user boredom or frustration, and provides an efficient, comfortable and targeted educational experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636398B_ABST
    Figure CN120636398B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent voice-assisted education system and method, relating to the technical field of assisted education. The system includes the following steps: firstly adjusting first voice content; calculating voice recognition accuracy and voice playback inclusiveness; updating the user's voice database; firstly marking key points in the voice; updating historical preference information; generating a knowledge list and corresponding knowledge exercises; and generating learning decisions based on playback scores. This invention improves user satisfaction by adjusting voice parameters according to the user's historical preferences, promotes in-depth learning by generating knowledge lists and exercises, enhances the flexibility and convenience of learning by sending the knowledge list to other devices, improves learning efficiency through voice highlighting and targeted practice, and avoids user boredom or frustration through playback scores and learning decisions, thus providing users with an efficient, comfortable, and targeted educational experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of auxiliary education, and in particular to an intelligent voice-assisted education method. Background Technology

[0002] In recent years, voice-guided educational products have leveraged large-scale modeling technology to engage in more natural dialogues and interactions with children. They employ advanced TTS technology and are deeply customized to children's language cognitive characteristics, optimizing speech rate, intonation, and vocabulary difficulty to make the voice more aligned with children's auditory habits and comprehension abilities. They also emphasize multimodal interaction design, combining various interaction methods such as voice, gestures, and facial expressions to provide children with a richer interactive experience.

[0003] Currently, Chinese invention patent CN113449068A discloses a voice interaction method and electronic device. This method receives first voice information from a second user via an electronic device, responds to the first voice information, identifies the first voice information, and uses the first voice information to request a voice dialogue with the first user. Based on the electronic device's recognition that the first voice information is the second user's voice information, the electronic device simulates the first user's voice and engages in a voice dialogue with the second user in the manner of a voice dialogue between the first and second users. However, this related technology lacks playback compensation for recognition vulnerabilities such as user-generated dialects, which can easily lead to incorrect playback. It also lacks compensation for playback parameters based on actual playback progress deviations, which is detrimental to playback accuracy. Furthermore, it does not make decisions regarding learning frequency based on educational outcomes, which is detrimental to the clarity of decision support, thus exhibiting certain limitations. Summary of the Invention

[0004] The technical problem solved by this invention is that related technologies do not compensate for playback errors due to recognition vulnerabilities such as the user's spoken dialect, which can easily lead to incorrect playback; they do not compensate for deviations in playback progress based on actual playback progress, which is detrimental to playback accuracy; and they do not make decisions on learning frequency based on educational outcomes, which is detrimental to the clarity of decision support, thus having certain limitations.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: Firstly, an intelligent voice-guided education method, comprising the following steps:

[0006] Step S100: Play the first voice content according to the user's instruction, and adjust the first voice parameters corresponding to the first voice content according to the user's historical preference information to obtain the second voice parameters, and play the first voice according to the second voice parameters;

[0007] Step S200: Obtain user feedback information after the first voice playback, calculate the voice recognition accuracy and voice playback inclusiveness based on the user feedback information, update the user voice database based on the voice recognition accuracy, mark the voice emphasis as a first mark, and update the historical preference information based on the voice playback inclusiveness.

[0008] Step S300: Generate a knowledge list and corresponding knowledge exercises based on the voice emphasis after the first mark, send the knowledge list to the linked device, set the user score of the knowledge exercises as the playback score, and generate a learning decision based on the playback score.

[0009] As a preferred embodiment of the intelligent voice-assisted education method of the present invention, the first voice content includes a voice name and a voice language, wherein the voice language is represented by a language and the dialect pronunciation corresponding to the language.

[0010] The historical preference information includes historical time period, historical timbre, historical volume, historical playback speed, and historical playback progress. The historical time period refers to the time period corresponding to the historical playback of the first audio content. The specific division logic of the historical time period includes:

[0011] Starting from 0:00, each 3-hour interval is a time period, until 24:00, completing the division of historical time periods;

[0012] The first voice parameters include timbre, volume, playback speed, and progress nodes. The progress node represents the end time point of the pulled progress bar. The first voice parameters are randomly set.

[0013] As a preferred embodiment of the intelligent voice-guided education method described in this invention, step S100 includes the following steps:

[0014] Step S101: Obtain user instructions, wherein the user instructions are voice instructions;

[0015] Step S102: According to the user's instruction, retrieve the user's corresponding voice language from the cloud database, and perform semantic recognition on the user's instruction based on the user's corresponding voice language. The semantic recognition is implemented based on a large language model, and the cloud database stores the user's voice language.

[0016] Step S103: Based on the semantics after semantic recognition, match the corresponding voice name and set the corresponding voice content as the first voice name;

[0017] Step S104: Retrieve the playback database, input the first voice name into the playback database, and retrieve the corresponding first voice segment.

[0018] In a preferred embodiment of the intelligent voice-guided education method described in this invention, the setting logic for the second voice parameter includes:

[0019] The system obtains the time point of receiving the user command, assigns the time point to a historical time period, and records it as the first time period. It retrieves historical preference information, filters the historical timbre, historical playback speed, and historical playback progress corresponding to the first time period, counts the first occurrence of each type of historical timbre, selects the historical timbre corresponding to the largest first occurrence, and sets the historical timbre corresponding to the largest first occurrence as the timbre of the second speech parameter.

[0020] Calculate the first average value of each historical volume, and set the first average value as the volume of the second voice parameter;

[0021] Calculate the second average value of each historical playback speed, and set the second average value as the playback speed of the second voice parameter;

[0022] Obtain the historical playback time period with the smallest time interval from the current time point, and set the end time point of the historical playback progress corresponding to the historical playback time period with the smallest time interval from the current time point as the progress node of the second voice parameter.

[0023] In a preferred embodiment of the intelligent voice-guided education method of the present invention, the first voice parameter is subjected to a first adjustment to obtain a second voice parameter, wherein the first adjustment includes:

[0024] The timbre in the first speech parameter is compared with the timbre in the second speech parameter. When the timbre in the first speech parameter is the same as the timbre in the second speech parameter, no first adjustment is performed. When the timbre in the first speech parameter is different from the timbre in the second speech parameter, the timbre in the first speech parameter is replaced with the timbre in the second speech parameter.

[0025] Adjust the playback speed and volume in the first voice parameter to be equal to the playback speed and volume in the second voice parameter, respectively;

[0026] Jump from the progress node of the first voice parameter to the progress node of the second voice parameter.

[0027] As a preferred embodiment of the intelligent voice-assisted education method of the present invention, the user feedback information includes switching voice content, looping voice content, and performing a second adjustment on the second voice parameter, wherein the second adjustment on the second voice parameter means performing a second adjustment on the progress node in the second voice parameter.

[0028] Obtain user feedback information after the first voice playback, calculate the voice recognition accuracy based on the switched voice content, and calculate the voice playback tolerance based on the looped voice content and the second adjustment of the second voice parameters.

[0029] The cloud database is updated based on the accuracy of speech recognition, the key points of the speech are marked first based on the inclusiveness of speech playback, and the historical preference information is updated.

[0030] The calculation logic for the speech recognition accuracy includes:

[0031] The system counts the second number of times the voice content is switched, counts the third number of times the user command is received, calculates the second ratio of the second number to the third number, and sets the second ratio as the voice recognition accuracy.

[0032] As a preferred embodiment of the intelligent voice-guided education method described in this invention, the calculation logic for the voice playback inclusiveness includes:

[0033] Obtain the first time interval between the progress node after the second adjustment and the original progress node in the second voice parameter; obtain the total duration of the voice segment; calculate the third ratio of the first time interval to the total duration of the voice segment; obtain the loop voice content; calculate the total duration of the loop voice content; calculate the fourth ratio of the total duration of the loop voice content to the total duration of the voice segment; calculate the second average of the third and fourth ratios; and set the second average as the voice playback tolerance.

[0034] As a preferred embodiment of the intelligent voice-assisted education method described in this invention, the cloud database is updated based on the accuracy of voice recognition, and the historical preference information is updated based on the inclusiveness of voice playback.

[0035] The logic for updating the cloud database based on speech recognition accuracy includes:

[0036] The first value is set as the speech recognition accuracy threshold. The speech recognition accuracy is compared with the first value. When the speech recognition accuracy is less than the first value, the speech of the current user command is set to the user's standard timbre, and the standard timbre replaces the user's original standard timbre in the cloud database.

[0037] The logic for updating historical preference information based on voice playback inclusiveness includes:

[0038] Obtain the time point corresponding to the progress node in the historical preference information, calculate the quotient between the time point corresponding to the progress node in the historical preference information and the voice playback tolerance, set the quotient as the time point corresponding to the progress node in the new historical preference information, and delete the time point corresponding to the progress node in the original historical preference information.

[0039] The first marking of speech focus includes marking the first time interval and the content of the loop speech, and automatically generating a knowledge list based on the language big model.

[0040] As a preferred embodiment of the intelligent voice-assisted education method described in this invention, the cloud database is updated based on the accuracy of voice recognition, and the historical preference information is updated based on the inclusiveness of voice playback.

[0041] The logic for updating the cloud database based on speech recognition accuracy includes:

[0042] The first value is set as the speech recognition accuracy threshold. The speech recognition accuracy is compared with the first value. When the speech recognition accuracy is less than the first value, the speech of the current user command is set to the user's standard timbre, and the standard timbre replaces the user's original standard timbre in the cloud database.

[0043] The logic for updating historical preference information based on voice playback inclusiveness includes:

[0044] Obtain the time point corresponding to the progress node in the historical preference information, calculate the quotient between the time point corresponding to the progress node in the historical preference information and the voice playback tolerance, set the quotient as the time point corresponding to the progress node in the new historical preference information, and delete the time point corresponding to the progress node in the original historical preference information.

[0045] The first marking of speech focus includes marking the first time interval and the content of the loop speech, and automatically generating a knowledge list based on the language big model.

[0046] Secondly, an intelligent voice-guided education system includes a control module, a computing module, and a decision-making module;

[0047] The control module plays the first voice content according to the user's instruction, and performs a first adjustment on the first voice parameters corresponding to the first voice content according to the user's historical preference information to obtain the second voice parameters, and plays the first voice according to the second voice parameters;

[0048] The calculation module obtains user feedback information after the first voice playback, calculates the voice recognition accuracy and voice playback inclusiveness based on the user feedback information, updates the user voice database based on the voice recognition accuracy, marks the voice emphasis as a first mark, and updates the historical preference information based on the voice playback inclusiveness.

[0049] The decision module generates a knowledge list and corresponding knowledge exercises based on the voice emphasis after the first mark, sends the knowledge list to the linked device, sets the user score of the knowledge exercises as the playback score, and generates a learning decision based on the playback score.

[0050] The beneficial effects of this invention are as follows: By adjusting voice parameters according to the user's historical preferences, the system can provide users with a voice playback experience that matches their personal habits, improving user satisfaction. Real-time updates to the voice recognition model and preference information based on user feedback continuously optimize the accuracy and comfort of voice interaction. By generating knowledge lists and exercises, the system helps users better understand and consolidate the knowledge points in the voice content. Learning decisions are generated based on the user's learning performance, providing targeted learning suggestions and promoting in-depth learning. Sending knowledge lists to other devices facilitates learning in different scenarios, enhancing the flexibility and convenience of learning. Through voice highlighting and targeted practice, the system helps users focus on key knowledge points, improving learning efficiency. By playing scores and making learning decisions, the system dynamically adjusts learning content according to the user's learning progress, preventing boredom or frustration. This intelligent voice-assisted education method combines personalization, dynamic optimization, and knowledge reinforcement, providing users with an efficient, comfortable, and targeted educational experience. Attached Figure Description

[0051] Figure 1 This is a schematic diagram of the basic process of an intelligent voice-guided education method provided in one embodiment of the present invention. Detailed Implementation

[0052] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0053] Example, refer to Figure 1 As an embodiment of the present invention, an intelligent voice-guided education method is provided, comprising the following steps:

[0054] Step S100: Play the first voice content according to the user's instruction, and adjust the first voice parameters corresponding to the first voice content according to the user's historical preference information to obtain the second voice parameters, and play the first voice according to the second voice parameters;

[0055] Step S200: Obtain user feedback information after the first voice playback, calculate the voice recognition accuracy and voice playback inclusiveness based on the user feedback information, update the user voice database based on the voice recognition accuracy, mark the voice emphasis as a first mark, and update the historical preference information based on the voice playback inclusiveness.

[0056] Step S300: Generate a knowledge list and corresponding knowledge exercises based on the voice emphasis after the first mark, send the knowledge list to the linked device, set the user score of the knowledge exercises as the playback score, and generate a learning decision based on the playback score.

[0057] This invention adjusts voice parameters based on the user's historical preferences, providing a personalized voice playback experience and enhancing user satisfaction. It continuously optimizes the accuracy and comfort of voice interaction by updating the voice recognition model and preference information in real time based on user feedback. By generating knowledge lists and exercises, the system helps users better understand and consolidate knowledge points within the voice content. It generates learning decisions based on the user's learning performance, providing targeted learning suggestions and promoting in-depth knowledge acquisition. The system can also send knowledge lists to other devices, facilitating learning in different scenarios and enhancing flexibility and convenience. Through voice highlighting and targeted practice, the system helps users focus on key knowledge points, improving learning efficiency. By playing scores and making learning decisions, the system dynamically adjusts learning content according to the user's learning progress, preventing boredom or frustration. This intelligent voice-assisted education method combines personalization, dynamic optimization, and knowledge reinforcement, providing users with an efficient, comfortable, and targeted educational experience.

[0058] The first audio content includes the audio name and the audio language, which is represented by the language and the corresponding dialect pronunciation.

[0059] Historical preference information includes historical time periods, historical timbre, historical volume, historical playback speed, and historical playback progress. The historical time period refers to the time period corresponding to the first audio content played in the past. The specific logic for dividing the historical time period includes:

[0060] Starting from 0:00, each 3-hour interval is a time period, until 24:00, completing the division of historical time periods;

[0061] The first voice parameters include timbre, volume, playback speed, and progress nodes. The progress node represents the end time point of the retrieved progress bar. The first voice parameters are randomly set.

[0062] In practical implementation, by supporting multiple languages ​​and dialects, the system can meet the needs of users from different regions and language backgrounds, improving the user experience. By dividing historical time periods in detail, the system can adjust voice parameters according to the user's preferences at different times. For example, users may prefer higher volume and faster playback speed in the morning, and lower volume and slower playback speed in the evening. Initially, voice parameters are randomly set to avoid monotony during first use. Subsequently, the system adjusts the parameters based on the user's historical preference information to ensure a more comfortable experience in subsequent uses. By recording progress nodes, the system can resume playback from where it left off, avoiding repeated listening to previously heard content. Adjusting timbre, volume, and playback speed based on the user's historical preferences provides a more personalized voice playback experience, thereby increasing user satisfaction and usage frequency. Supporting voice content in multiple languages ​​and dialects not only meets the language needs of different users but also serves as an auxiliary tool for language learning, helping users learn and practice the pronunciation of different languages ​​and dialects.

[0063] Step S100 includes the following steps:

[0064] Step S101: Obtain user instructions, which are represented as voice instructions;

[0065] Step S102: According to the user's instruction, retrieve the user's corresponding voice language from the cloud database, and perform semantic recognition on the user's instruction based on the user's corresponding voice language. The semantic recognition is implemented based on the language model, and the cloud database stores the user's voice language.

[0066] Step S103: Based on the semantics after semantic recognition, match the corresponding voice name and set the corresponding voice content as the first voice name;

[0067] Step S104: Retrieve the playback database, input the first voice name into the playback database, and retrieve the corresponding first voice segment.

[0068] In practice, by retrieving the user's corresponding voice language from the cloud database and using a large language model for semantic recognition, the system can more accurately understand the user's instructions and avoid misunderstandings caused by language differences. The semantic recognition technology based on the large language model can better understand the user's intentions, accurately parsing even complex expressions or dialect accents in the instructions. Based on the user's specified voice language and voice name, it accurately matches and plays the content the user wants, avoiding the trouble of manually searching through a large number of voice resources. The cloud database stores the user's voice language preferences and a large amount of voice content, enabling the system to flexibly adapt to the personalized needs of different users. As the user's language preferences change or the voice content is updated, the system dynamically adjusts the semantic recognition model and playback database to maintain the system's efficiency and accuracy.

[0069] The logic for setting the second voice parameter includes:

[0070] Get the time point of receiving the user command, classify the time point of receiving the command into a historical time period, and mark it as the first time period. Retrieve historical preference information, filter the historical timbre, historical playback speed and historical playback progress corresponding to the first time period, count the first occurrence of each type of historical timbre, select the historical timbre corresponding to the largest first occurrence, and set the historical timbre corresponding to the largest first occurrence as the timbre of the second speech parameter.

[0071] Calculate the first average value of each historical volume, and set the first average value as the volume of the second voice parameter;

[0072] Calculate the second average value of each historical playback speed, and set the second average value as the playback speed of the second voice parameter;

[0073] Obtain the historical playback time segment with the smallest time interval from the current time point, and set the end time point of the historical playback progress corresponding to the historical playback time segment with the smallest time interval from the current time point as the progress node of the second voice parameter.

[0074] In practice, by assigning the time of user commands to a specific time period and retrieving historical preference information within that period, the system can accurately match the user's preferences within that specific time frame. For example, a user might prefer a lively tone in the morning and a gentle tone in the evening. The settings for tone, volume, and playback speed are based on the user's historical preference data within that specific time period, providing a voice playback experience that best suits their current habits. The system automatically counts the frequency of historical tone occurrences and calculates the average volume and playback speed without manual intervention, enhancing the system's intelligence. As user habits change, the system dynamically updates historical preference information and adjusts voice parameters in real time to ensure the user experience is always optimal. By setting the progress node to the end of the last playback, the system can achieve seamless playback, preventing users from repeatedly listening to previously heard content and improving ease of use. This not only enhances the user experience but also strengthens the system's intelligence, better meeting users' personalized needs and thus increasing user satisfaction and usage frequency.

[0075] The first speech parameter is modulated in a first manner to obtain the second speech parameter. The first modulation includes:

[0076] The timbre in the first speech parameter is compared with the timbre in the second speech parameter. When the timbre in the first speech parameter is the same as the timbre in the second speech parameter, no first adjustment is performed. When the timbre in the first speech parameter is different from the timbre in the second speech parameter, the timbre in the first speech parameter is replaced with the timbre in the second speech parameter.

[0077] Adjust the playback speed and volume in the first voice parameter to be equal to the playback speed and volume in the second voice parameter, respectively;

[0078] Jump from the progress node of the first voice parameter to the progress node of the second voice parameter.

[0079] User feedback information includes switching voice content, looping voice content, and performing a second adjustment on the second voice parameter. Performing a second adjustment on the second voice parameter means performing a second adjustment on the progress node in the second voice parameter.

[0080] Obtain user feedback information after the first voice playback, calculate the voice recognition accuracy based on the switched voice content, and calculate the voice playback tolerance based on the looped voice content and the second adjustment of the second voice parameters.

[0081] The cloud database is updated based on the accuracy of speech recognition, the key points of the speech are marked first based on the inclusiveness of speech playback, and the historical preference information is updated.

[0082] The calculation logic for speech recognition accuracy includes:

[0083] The system counts the second number of times the voice content is switched, counts the third number of times the user command is received, calculates the second ratio of the second and third counts, and sets the second ratio as the voice recognition accuracy.

[0084] In practice, by statistically analyzing the frequency of user switching of voice content and the total number of commands, the system can quantify the accuracy of voice recognition and update the voice recognition model accordingly. By analyzing the content that users loop and the behavior of adjusting progress nodes, the system can identify the key content that users are interested in and mark it as voice focus. Based on the voice playback inclusiveness, the system can dynamically adjust the voice playback parameters to better meet the user's preferences and needs. By optimizing voice recognition and playback parameters, the system can reduce the situation where users frequently intervene due to dissatisfaction (such as switching voice content or adjusting progress nodes), which not only improves the utilization efficiency of educational resources but also enhances the user's learning effect.

[0085] The calculation logic for voice playback inclusiveness includes:

[0086] Obtain the first time interval between the progress node after the second adjustment and the original progress node in the second voice parameter; obtain the total duration of the voice segment; calculate the third ratio of the first time interval to the total duration of the voice segment; obtain the loop voice content; calculate the total duration of the loop voice content; calculate the fourth ratio of the total duration of the loop voice content to the total duration of the voice segment; calculate the second average of the third and fourth ratios; and set the second average as the voice playback tolerance.

[0087] The cloud database is updated based on the accuracy of speech recognition, and historical preference information is updated based on the inclusiveness of speech playback.

[0088] The logic for updating the cloud database based on speech recognition accuracy includes:

[0089] Set the first value as the speech recognition accuracy threshold, compare the speech recognition accuracy with the first value, and when the speech recognition accuracy is less than the first value, set the speech of the current user command to the user's standard voice, and replace the user's original standard voice in the cloud database with the standard voice.

[0090] The logic for updating historical preference information based on voice playback inclusiveness includes:

[0091] Obtain the time point corresponding to the progress node in the historical preference information, calculate the quotient between the time point corresponding to the progress node in the historical preference information and the voice playback tolerance, set the quotient as the time point corresponding to the progress node in the new historical preference information, and delete the time point corresponding to the progress node in the original historical preference information.

[0092] The first marking of speech focus includes marking the first time interval and the content of the loop speech, and automatically generating a knowledge list based on the language big model.

[0093] In practice, by comparing speech recognition accuracy with a threshold and updating the user's standard timbre when the accuracy is low, the system can better adapt to the user's pronunciation characteristics, reducing misunderstandings and misrecognitions. By calculating the quotient and updating the progress nodes, the system can dynamically reflect changes in user preferences over different time periods. This dynamic adjustment mechanism makes the system's personalized settings more flexible and better meets the user's real-time needs. Deleting old progress node time points avoids excessive redundancy in historical preference information and improves the system's operating efficiency. By marking the content played in loops and the content played in the first time interval, the system can accurately identify the key content that the user is interested in. Combined with the language big data model, it can automatically generate a knowledge list, providing richer and more accurate learning content.

[0094] Generate a knowledge list from the key points of the audio after the first mark;

[0095] Randomly delete words from the knowledge list to create fill-in-the-blank questions, and then set these fill-in-the-blank questions as corresponding knowledge exercises.

[0096] Send the knowledge list to the linked device, which is the user's learning device;

[0097] Set the user's score for the knowledge exercises as the playback score.

[0098] The logic for generating learning decisions based on playback scores includes:

[0099] Set the first score as the score threshold, compare the playback score with the first score, and when the playback score is less than the first score, set the learning cycle of the voice focus corresponding to the knowledge list as the first cycle. When the playback score is greater than or equal to the first score, set the learning cycle of the voice focus corresponding to the knowledge list as the second cycle. The first cycle represents a duration that is less than the duration represented by the second cycle.

[0100] In practice, by generating knowledge lists and exercises, the system helps users systematically review and consolidate key knowledge points from the audio content. The fill-in-the-blank format tests users' grasp of the details of knowledge points, helping them to better understand and memorize. The system dynamically adjusts the learning cycle based on user scores, providing users with more targeted learning plans to help them review efficiently. The knowledge list is sent to the user's connected devices, allowing users to learn anytime, anywhere on different devices, enhancing the flexibility and convenience of learning. The dynamic adjustment of the learning cycle ensures that users review key knowledge within an appropriate timeframe, improving learning efficiency.

[0101] This invention adjusts voice parameters based on the user's historical preferences, providing a personalized voice playback experience and enhancing user satisfaction. It continuously optimizes the accuracy and comfort of voice interaction by updating the voice recognition model and preference information in real time based on user feedback. By generating knowledge lists and exercises, the system helps users better understand and consolidate knowledge points within the voice content. It generates learning decisions based on the user's learning performance, providing targeted learning suggestions and promoting in-depth knowledge acquisition. The system can also send knowledge lists to other devices, facilitating learning in different scenarios and enhancing flexibility and convenience. Through voice highlighting and targeted practice, the system helps users focus on key knowledge points, improving learning efficiency. By playing scores and making learning decisions, the system dynamically adjusts learning content according to the user's learning progress, preventing boredom or frustration. This intelligent voice-assisted education method combines personalization, dynamic optimization, and knowledge reinforcement, providing users with an efficient, comfortable, and targeted educational experience.

[0102] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium is implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0103] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention should all be covered within the scope of the claims of the present invention.

Claims

1. A method for intelligent voice-guided education, characterized in that, Includes the following steps: Step S100: Play the first voice content according to the user's instruction, and adjust the first voice parameters corresponding to the first voice content according to the user's historical preference information to obtain the second voice parameters, and play the first voice according to the second voice parameters; Step S200: Obtain user feedback information after the first voice playback, calculate the voice recognition accuracy and voice playback inclusiveness based on the user feedback information, update the user voice database based on the voice recognition accuracy, mark the voice emphasis as a first mark, and update the historical preference information based on the voice playback inclusiveness. The user feedback information includes switching voice content, looping voice content, and performing a second adjustment on the second voice parameter. The second adjustment on the second voice parameter means performing a second adjustment on the progress node in the second voice parameter. Obtain user feedback information after the first voice playback, calculate the voice recognition accuracy based on the switched voice content, and calculate the voice playback tolerance based on the looped voice content and the second adjustment of the second voice parameters. The cloud database is updated based on the accuracy of speech recognition, the key points of the speech are marked first based on the inclusiveness of speech playback, and the historical preference information is updated. The calculation logic for the speech recognition accuracy includes: The second number of times the voice content was switched is counted, the third number of times the user command was received is counted, the second ratio of the second number to the third number is calculated, and the second ratio is set as the voice recognition accuracy. The calculation logic for the voice playback inclusiveness includes: Obtain the first time interval between the progress node after the second adjustment and the original progress node in the second voice parameter; obtain the total duration of the voice segment; calculate the third ratio of the first time interval to the total duration of the voice segment; obtain the loop voice content; calculate the total duration of the loop voice content; calculate the fourth ratio of the total duration of the loop voice content to the total duration of the voice segment; calculate the second average of the third ratio and the fourth ratio; and set the second average as the voice playback tolerance. The first marking of speech focus includes marking the first time interval and the content of cyclic speech, and automatically generating a knowledge list based on the language big model; Step S300: Generate a knowledge list and corresponding knowledge exercises based on the voice emphasis after the first mark, send the knowledge list to the linked device, set the user score of the knowledge exercises as the playback score, and generate a learning decision based on the playback score.

2. The intelligent voice-guided education method as described in claim 1, characterized in that: The first audio content includes an audio name and an audio language, wherein the audio language is represented by a language and the corresponding dialect pronunciation; The historical preference information includes historical time period, historical timbre, historical volume, historical playback speed, and historical playback progress. The historical time period refers to the time period corresponding to the historical playback of the first audio content. The specific division logic of the historical time period includes: Starting from 0:00, each 3-hour interval is a time period, until 24:00, completing the division of historical time periods; The first voice parameters include timbre, volume, playback speed, and progress nodes. The progress node represents the end time point of the pulled progress bar. The first voice parameters are randomly set.

3. The intelligent voice-guided education method as described in claim 1, characterized in that: Step S100 includes the following steps: Step S101: Obtain user instructions, wherein the user instructions are voice instructions; Step S102: According to the user's instruction, retrieve the user's corresponding voice language from the cloud database, and perform semantic recognition on the user's instruction based on the user's corresponding voice language. The semantic recognition is implemented based on a large language model, and the cloud database stores the user's voice language. Step S103: Based on the semantics after semantic recognition, match the corresponding voice name and set the corresponding voice content as the first voice name; Step S104: Retrieve the playback database, input the first voice name into the playback database, and retrieve the corresponding first voice segment.

4. The intelligent voice-guided education method as described in claim 3, characterized in that: The logic for setting the second voice parameter includes: The receiving time point of the user command is obtained, and the receiving time point is assigned to a historical time period, which is recorded as the first time period. Historical preference information is retrieved, and the historical timbre, historical playback speed and historical playback progress corresponding to the first time period are filtered. The first occurrence number of each type of historical timbre is counted, and the historical timbre corresponding to the largest first occurrence number is selected. The historical timbre corresponding to the largest first occurrence number is set as the timbre of the second voice parameter. Calculate the first average value of each historical volume, and set the first average value as the volume of the second voice parameter; Calculate the second average value of each historical playback speed, and set the second average value as the playback speed of the second voice parameter; Obtain the historical playback time period with the smallest time interval from the current time point, and set the end time point of the historical playback progress corresponding to the historical playback time period with the smallest time interval from the current time point as the progress node of the second voice parameter.

5. The intelligent voice-guided education method as described in claim 4, characterized in that: The first speech parameter is subjected to a first modulation to obtain the second speech parameter, wherein the first modulation includes: The timbre in the first speech parameter is compared with the timbre in the second speech parameter. When the timbre in the first speech parameter is the same as the timbre in the second speech parameter, no first adjustment is performed. When the timbre in the first speech parameter is different from the timbre in the second speech parameter, the timbre in the first speech parameter is replaced with the timbre in the second speech parameter. Adjust the playback speed and volume in the first voice parameter to be equal to the playback speed and volume in the second voice parameter, respectively; Jump from the progress node of the first voice parameter to the progress node of the second voice parameter.

6. The intelligent voice-guided education method as described in claim 1, characterized in that: The cloud database is updated based on the accuracy of speech recognition, and historical preference information is updated based on the inclusiveness of speech playback. The logic for updating the cloud database based on speech recognition accuracy includes: The first value is set as the speech recognition accuracy threshold. The speech recognition accuracy is compared with the first value. When the speech recognition accuracy is less than the first value, the speech of the current user command is set to the user's standard timbre, and the standard timbre replaces the user's original standard timbre in the cloud database. The logic for updating historical preference information based on voice playback inclusiveness includes: Obtain the time point corresponding to the progress node in the historical preference information, calculate the quotient between the time point corresponding to the progress node in the historical preference information and the voice playback tolerance, set the quotient as the time point corresponding to the progress node in the new historical preference information, and delete the time point corresponding to the progress node in the original historical preference information.

7. The intelligent voice-guided education method as described in claim 6, characterized in that: Generate a knowledge list from the key points of the audio after the first mark; Randomly delete words from the knowledge list to create fill-in-the-blank questions, and set these fill-in-the-blank questions as corresponding knowledge exercises; The knowledge list is sent to the linked device, which refers to the user's learning device; Set the user's score for the knowledge exercises as the playback score. The logic for generating learning decisions based on playback scores includes: Set the first score as the score threshold, compare the playback score with the first score, and when the playback score is less than the first score, set the learning cycle of the voice focus corresponding to the knowledge list as the first cycle. When the playback score is greater than or equal to the first score, set the learning cycle of the voice focus corresponding to the knowledge list as the second cycle. The first cycle represents a duration that is less than the duration represented by the second cycle.

8. An intelligent voice-guided education system, wherein the system is used to execute the intelligent voice-guided education method according to claim 1, characterized in that, It includes a control module, a calculation module, and a decision-making module; The control module plays the first voice content according to the user's instruction, and performs a first adjustment on the first voice parameters corresponding to the first voice content according to the user's historical preference information to obtain the second voice parameters, and plays the first voice according to the second voice parameters; The calculation module obtains user feedback information after the first voice playback, calculates the voice recognition accuracy and voice playback inclusiveness based on the user feedback information, updates the user voice database based on the voice recognition accuracy, marks the voice emphasis as a first mark, and updates the historical preference information based on the voice playback inclusiveness. The decision module generates a knowledge list and corresponding knowledge exercises based on the voice emphasis after the first mark, sends the knowledge list to the linked device, sets the user score of the knowledge exercises as the playback score, and generates a learning decision based on the playback score.

Citation Information

Patent Citations

  • Voice interaction method and electronic equipment

    CN113449068A

  • Text-based speech playing method, device, computer device and storage medium

    CN108109636A

  • Voice detector

    JP2000057325A