3D virtual image hearing and speech training method based on AI interaction
By using AI-interactive 3D virtual avatar auditory speech training methods, combined with intonation auditory methods and 3D virtual avatars, the problem of insufficient pronunciation teaching in English teaching has been solved, improving students' pronunciation level and learning outcomes, and achieving educational equity and balanced development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUNNAN BEIFEI TECH CO LTD
- Filing Date
- 2022-09-09
- Publication Date
- 2026-04-24
AI Technical Summary
In English teaching, teachers' neglect of pronunciation instruction leads to students' non-standard pronunciation, which affects their subsequent learning. Existing technologies are difficult to effectively improve students' pronunciation skills and lack timeliness and fairness.
This method employs an AI-based interactive 3D virtual avatar auditory speech training approach, combining intonation auditory methods and 3D virtual avatars. It uses low-pass filtering to process speech signals, combined with body movement stimulation and oral practice, and utilizes AI speech recognition and evaluation to provide a brand-new method for learning English speech.
It effectively improves English learners' pronunciation, compensates for deficiencies in classroom teaching, achieves educational equity and balanced development, enhances learners' motor skills, spatial orientation and memory span, and provides immediate feedback and standardized assessment.
Smart Images

Figure CN115910036B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of modern educational technology, and in particular to a method for training auditory and speech perception using 3D virtual avatars based on AI interaction. Background Technology
[0002] Speech enables speakers to express their thoughts, and correct pronunciation makes it easier for listeners to understand. Therefore, the first step in learning English is learning correct pronunciation. Numerous studies have shown that the effectiveness of pronunciation learning in the early stages of English learning directly affects the entire subsequent English learning process. However, in actual English teaching, teachers often focus on vocabulary and grammar instruction, neglecting pronunciation teaching and practice. This is especially true when English teachers themselves lack good pronunciation skills and experience in teaching pronunciation, leading to users' incorrect pronunciation and hindering their subsequent learning. Summary of the Invention
[0003] The purpose of this invention is to at least address the problems existing in the current education process and provide an AI-interactive 3D virtual image auditory speech training method. This new English pronunciation learning method combines auditory training with physical stimulation and oral (speaking) practice. It combines intonation auditory method with 3D virtual image, AI speech recognition and evaluation, which can effectively improve the pronunciation level of English learners, make up for the deficiencies in pronunciation teaching in the classroom and the inadequacy of teachers' varying pronunciation conditions, and solve the problem of timeliness in classroom teaching. More importantly, such a method that can be operated on mobile smart terminals is a positive innovation to promote educational equity and achieve common and balanced educational development.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] A method for training auditory and speech perception of 3D virtual avatars based on AI interaction, characterized by the following steps:
[0006] S1. The recorded speech signal undergoes low-pass filtering to convert it into a low-frequency speech pattern, retaining low frequencies below 300Hz. These low-frequency frequencies preserve the prosodic features of the speech, including stress, rhythm, loudness, and intonation. This process removes high-frequency words that are essential for word recognition, while the low-frequency speech signal, which retains the prosodic features, effectively reduces the learning learner's semantic and syntactic processing load, freeing up more attentional resources for other cognitive processes.
[0007] S2. The processed input speech signals are divided into eight sentence types: affirmative statements, negative statements, general questions, special questions, alternative questions, tag questions, imperative sentences, and exclamatory sentences. A graded sentence library is obtained, and users are continuously trained for at least 30 seconds. The audio of the sentences used for training undergoes low-pass filtering to highlight speech parameters such as rhythm, intonation, pitch, tension, pauses, duration, and loudness, which can enhance the user's proprioceptive perception of the language signals. In this training, through sensory integration training involving senses, hearing, and vision, learners can maximize the development of neural pathways in the brain and expand their learning potential. Specifically, users coordinate their body movements with the rhythm of the speech signals to improve their motor skills, spatial orientation, and memory span, achieving coordinated development of proprioception, hearing, and speech.
[0008] S3. Create a virtual image. Create a virtual character in Unity 3D and match the skeletal structure of the virtual character with the predefined skeletal structure of Mecanim. Create an Animator component for each action to build animation effects such as waving arms, rotating, and beating. Each action is triggered by a controller instruction. For each action effect, corresponding parameters need to be set to control the amplitude of the action.
[0009] S4. Integrate the graded sentence library with the corresponding 3D virtual imaging. The 3D virtual character's movements have two forms: one is the up-and-down movement of the arms, and the other is the opening and closing of the hands. By obtaining the highest and lowest points of the audio fundamental frequency, set the highest point of the arm waving or the maximum degree of the hands opening in the 3D virtual character as the highest point of the audio fundamental frequency; set the lowest point of the arm waving or the minimum degree of the hands closing in the 3D virtual character as the lowest point of the audio fundamental frequency. The amplitude of the 3D virtual character at the highest and lowest points of the audio fundamental frequency is 100% * (current frequency – minimum frequency) / (maximum frequency – minimum frequency) = amplitude % (amplitude is between 0 and 100).
[0010] During the actual animation playback, users can choose between two animation formats, or the system can randomly display one. The system obtains the current time in real time, and the 3D virtual character displays the animation based on the time and the calculated movement amplitude. When an undefined value is encountered, the animation does not change until a valid frequency parameter is obtained.
[0011] S5. Playing the speech signal as a unit, the low-frequency speech of 0Hz to 300Hz in the two channels is separated by a frequency band separator. This includes the fundamental frequency (F0) that determines the pitch. The fundamental frequency determines the pitch change of the intonation. Based on the fundamental frequency curve of each sentence, this method sets the animation limit value and motion trajectory of the 3D image to guide the user to make appropriate body rhythms with the intonation of the sentence.
[0012] S6. The words in the graded sentence library are divided into high and low pitch words and displayed as word levels. While listening to the low-pass filtered audio material, the user watches a 3D virtual image with body movements. In this step, the user needs to imitate the low-pass filtered speech and make body movements in reference to the 3D virtual image. The 3D virtual image will use the rhythm of the sentence as the melody, and the cartoon character will make appropriate body movements in accordance with the stress, rhythm and intonation of the sentence, such as moving the arms up and down or opening and closing the hands in accordance with the changes in the pitch and intonation of the sentence. In this step, the user imitates the low-pass filtered speech while making body movements.
[0013] S7. AI Assessment: The AI assessment module assesses the user's pronunciation. The assessment module includes a display module for intonation, actions, and their combinations; a module for marking intonation, actions, and their combinations to form an assessment; a module for scoring the marked results; a result output module; and a module for assessment results and suggestions.
[0014] Preferably, it includes a data acquisition module for acquiring voice data and motion data from multiple users; the voice data is collected by a voice acquisition device and includes a 3D virtual image seen from the user's perspective; the motion data is collected by a motion acquisition device and includes vector data generated by the user's motion.
[0015] The scenario generation module is used to select and generate training scenarios from multiple preset training scenarios in response to user actions.
[0016] The role matching module is used to determine the content that the user sees in the training scenario based on the user's voice data, and to match the user's corresponding hand movements from the training scenario based on the content seen.
[0017] The motion generation module is used to generate the motion corresponding to the character in the training scenario based on the user's motion data.
[0018] The panoramic integration module is used to integrate the roles and actions of the multiple users and output them in panoramic mode;
[0019] The output in panoramic mode includes: when performing panoramic integration, determining the aspect ratio of the panoramic mode based on the degree of user dispersion; or when performing panoramic integration, extracting videos containing users from the scene and stitching the videos containing users together.
[0020] The feedback module is used to determine whether the difference between the action and the reference action is within a preset range, and when it is not within the preset range, it generates feedback information; the feedback information includes reminding the user of the current improper operation through vibration, specifically: using a patch on the part corresponding to the action that is not within the preset range to remind the user of the improper movement.
[0021] Preferably, the data acquisition module includes:
[0022] The identification unit is used to generate a corresponding user identification number for each user.
[0023] The matching unit is used to match the user identifier with the corresponding eye-tracking acquisition device and motion acquisition device;
[0024] The acquisition unit is used to generate eye movement data and motion data for each user through the eye movement acquisition device and the motion acquisition device.
[0025] Preferably, the data acquisition module is further configured to acquire audio data from multiple users; the panoramic integration module is further configured to integrate the audio data with the user's corresponding role; the interactive training system further includes:
[0026] An audio playback module is used to receive hand gestures from the 3D virtual avatar or the user and play the corresponding audio data.
[0027] Preferably, the data acquisition module is further configured to acquire audio data from multiple users;
[0028] The audio adjustment module is used to integrate the roles, actions, and audio data of the multiple users, and adjust the volume according to the distance of the roles in the panoramic mode and store the data.
[0029] Preferably, the stored data is tested and scored using AI evaluation.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] (1) This invention discloses a novel English pronunciation learning method that combines intonation auditory method with 3D virtual image, AI speech recognition and evaluation. It is easy to operate and can effectively improve the pronunciation level of English learners, make up for the deficiencies in pronunciation teaching in the classroom and the lack of teachers' varying pronunciation conditions, solve the problem of timeliness in classroom teaching. More importantly, such a method that can be operated on a mobile smart terminal is a positive innovation to promote educational equity and achieve common and balanced development of education. Attached Figure Description
[0032] The present invention will be further described below with reference to the accompanying drawings and embodiments:
[0033] Figure 1 Low-pass filtered audio spectrogram;
[0034] Figure 2 The audio spectrogram is not filtered.
[0035] Figure 3 Schematic diagram of declarative sentence patterns;
[0036] Figure 4 This is a schematic diagram of a general question pattern;
[0037] Figure 5 This is a schematic diagram of a special interrogative sentence.
[0038] Figure 6 This is a diagram of an imperative sentence;
[0039] Figure 7 This is a schematic diagram of a negative statement.
[0040] Figure 8 This is a schematic diagram of tag questions;
[0041] Figure 9 A schematic diagram of the choice question pattern;
[0042] Figure 10 This is a diagram of an exclamation mark. Detailed Implementation
[0043] This invention provides a technical solution: a 3D virtual image auditory speech training method based on AI interaction, characterized by the following steps:
[0044] S1. The recorded speech signal undergoes low-pass filtering to convert it into a low-frequency speech pattern, retaining low frequencies below 300Hz. This low-frequency audio preserves the prosodic features of the speech, including stress, rhythm, loudness, and intonation. In this way, high-frequency frequencies that enable word recognition are removed, while the low-frequency speech signal, preserving the prosodic features, effectively reduces the learning learner's semantic and syntactic processing load, freeing up more attentional resources for other cognitive processes.
[0045] S2. The processed input speech signals are divided into eight sentence types: affirmative statements, negative statements, general questions, special questions, alternative questions, tag questions, imperative sentences, and exclamatory sentences. A graded sentence library is obtained, and users are continuously trained for at least 30 seconds. The audio of the sentences used for training undergoes low-pass filtering to highlight speech parameters such as rhythm, intonation, pitch, tension, pauses, duration, and loudness, which can enhance the user's proprioceptive perception of the language signals. In this training, through sensory integration training involving senses, hearing, and vision, learners can maximize the development of neural pathways in the brain and expand their learning potential. Specifically, users coordinate their body movements with the rhythm of the speech signals to improve their motor skills, spatial orientation, and memory span, achieving coordinated development of proprioception, hearing, and speech.
[0046] S3. Create a virtual image. Create a virtual character in Unity 3D and match the skeletal structure of the virtual character with the predefined skeletal structure of Mecanim. Create an Animator component for each action to build animation effects such as waving arms, rotating, and beating. Each action is triggered by a controller instruction. For each action effect, corresponding parameters need to be set to control the amplitude of the action.
[0047] S4. Integrate the graded sentence library with the corresponding 3D virtual imaging. The 3D virtual character's movements have two forms: one is the up-and-down movement of the arms, and the other is the opening and closing of the hands. By obtaining the highest and lowest points of the audio fundamental frequency, set the highest point of the arm waving or the maximum degree of the hands opening in the 3D virtual character as the highest point of the audio fundamental frequency; set the lowest point of the arm waving or the minimum degree of the hands closing in the 3D virtual character as the lowest point of the audio fundamental frequency. The amplitude of the 3D virtual character at the highest and lowest points of the audio fundamental frequency is 100% * (current frequency – minimum frequency) / (maximum frequency – minimum frequency) = amplitude % (amplitude is between 0 and 100).
[0048] While the audio is playing, the current time is obtained in real time, and the 3D animation is displayed based on the time and the calculated range of motion.
[0049] When an undefined value is encountered, the animation does not change until a valid frequency parameter is obtained.
[0050] As shown in the table below:
[0051]
[0052] During the actual animation playback, users can choose between two animation formats, or the system can randomly display one. The system obtains the current time in real time, and the 3D virtual character displays the animation based on the time and the calculated movement amplitude. When an undefined value is encountered, the animation does not change until a valid frequency parameter is obtained.
[0053] S5. Playing the speech signal as a unit, the low-frequency speech of 0Hz to 300Hz in the two channels is separated by a frequency band separator. This includes the fundamental frequency (F0) that determines the pitch. The fundamental frequency determines the pitch change of the intonation. Based on the fundamental frequency curve of each sentence, this method sets the animation limit value and motion trajectory of the 3D image to guide the user to make appropriate body rhythms with the intonation of the sentence.
[0054] S6. The words in the graded sentence library are divided into high and low pitch words and displayed as word levels. While listening to the low-pass filtered audio material, the user watches a 3D virtual image with body movements. In this step, the user needs to imitate the low-pass filtered speech and make body movements in reference to the 3D virtual image. The 3D virtual image will use the rhythm of the sentence as the melody, and the cartoon character will make appropriate body movements in accordance with the stress, rhythm and intonation of the sentence, such as moving the arms up and down or opening and closing the hands in accordance with the changes in the pitch and intonation of the sentence. In this step, the user imitates the low-pass filtered speech while making body movements.
[0055] S7. AI Assessment: The AI assessment module assesses the user's pronunciation. The assessment module includes a display module for intonation, actions, and their combinations; a module for marking intonation, actions, and their combinations to form an assessment; a module for scoring the marked results; a result output module; and a module for assessment results and suggestions.
[0056] Preferably, it includes a data acquisition module for acquiring voice data and motion data from multiple users; the voice data is collected by a voice acquisition device and includes a 3D virtual image seen from the user's perspective; the motion data is collected by a motion acquisition device and includes vector data generated by the user's motion.
[0057] The scenario generation module is used to select and generate training scenarios from multiple preset training scenarios in response to user actions.
[0058] The role matching module is used to determine the content that the user sees in the training scenario based on the user's voice data, and to match the user's corresponding hand movements from the training scenario based on the content seen.
[0059] The motion generation module is used to generate the motion corresponding to the character in the training scenario based on the user's motion data.
[0060] The panoramic integration module is used to integrate the roles and actions of the multiple users and output them in panoramic mode;
[0061] The output in panoramic mode includes: when performing panoramic integration, determining the aspect ratio of the panoramic mode based on the degree of user dispersion; or when performing panoramic integration, extracting videos containing users from the scene and stitching the videos containing users together.
[0062] The feedback module is used to determine whether the difference between the action and the reference action is within a preset range, and when it is not within the preset range, it generates feedback information; the feedback information includes reminding the user of the current improper operation through vibration, specifically: using a patch on the part corresponding to the action that is not within the preset range to remind the user of the improper movement.
[0063] Preferably, the data acquisition module includes:
[0064] The identification unit is used to generate a corresponding user identification number for each user.
[0065] The matching unit is used to match the user identifier with the corresponding eye-tracking acquisition device and motion acquisition device;
[0066] The acquisition unit is used to generate eye movement data and motion data for each user through the eye movement acquisition device and the motion acquisition device.
[0067] Preferably, the data acquisition module is further configured to acquire audio data from multiple users; the panoramic integration module is further configured to integrate the audio data with the user's corresponding role; the interactive training system further includes:
[0068] An audio playback module is used to receive hand gestures from the 3D virtual avatar or the user and play the corresponding audio data.
[0069] Preferably, the data acquisition module is further configured to acquire audio data from multiple users;
[0070] The audio adjustment module is used to integrate the roles, actions, and audio data of the multiple users, and adjust the volume according to the distance of the roles in the panoramic mode and store the data.
[0071] Preferably, the stored data is tested and scored using AI evaluation.
[0072] This invention's training method is based on the intonation-based auditory method, an auditory-centric approach that develops spoken language through binaural listening. The intonation-based auditory method is grounded in the brain's neuroplasticity, using sensory integration training involving sound, hearing, and vision to maximize the development of neural pathways in the brain and expand learning potential. This multi-sensory integration connects the learner's auditory cortex with related vestibular and motor areas (especially speech areas), linking the brain, body, and vocal organs. It combines auditory training with physical stimulation (vestibular system) and oral (speaking) practice, reorganizing and strengthening neural pathway connections to improve learning outcomes. Figure 1As shown, this method views speech perception and oral production as a multi-sensory, holistic experience, allowing the vestibular system, motor skills, and vocalization to develop simultaneously. The aim is to maximize the use of neuroplasticity to reorganize neural pathways in the brain, thereby improving auditory and oral language skills. The intonation-based auditory method has first achieved good results in clinical practice, improving the auditory and speaking skills of children and adults with hearing loss; simultaneously, it can significantly improve the foreign language proficiency of foreign language learners. Studies on French, Chinese, and English learners have found that the intonation-based auditory method effectively improves French phonemes, Chinese tones, comprehensive English oral skills, English listening comprehension, pronunciation correction, and phonological working memory.
[0073] The effectiveness of the intonation-based auditory method in speech therapy and foreign language learning lies in its emphasis on rhythm and speech patterns, as these are the foundation of listening and speaking skills. Through low-frequency speech patterns, the vestibular system and cochlea are stimulated by variations in rhythm and intonation, and these two organs are particularly sensitive to the rhythm of speech. This is because the development of the auditory organs begins with the perception of low-frequency sounds (speech rhythm) in the womb. Infants begin to develop proprioceptive memory, laying the foundation for later auditory memory development—that is, from "feeling" to "hearing." In speech, the vestibular and auditory systems are more sensitive to prosodic signals below 300Hz. Therefore, passing the speech signal through a low-pass filter to retain low frequencies below 300Hz preserves the prosodic features of speech, including stress, rhythm, loudness, and intonation. In this way, high-frequency words that can be recognized are removed, while low-frequency speech signals that retain the rhythm of speech effectively reduce the processing load on learners' semantic and syntactic processing, freeing up more attentional resources for other cognitive processes. The main functions of the vestibular system are to perceive body movements and gravity. In the peripheral system, the vestibular system is part of the inner ear, connecting to the cochlea. Hearing develops from vestibular perception, and the two complement each other, simultaneously sensing and hearing speech. The importance of the vestibular system lies in its role as the integrative and organizing system of all senses, generating spatial perception. All sensory input needs to be integrated through the body, which is why vestibular perception needs to be trained and stimulated. Only when vestibular perception and body movement are unified and coordinated in vocalization can neurons maximize the development of new synapses connecting other neurons due to neuroplasticity. Speech information is received through proprioception and vestibular terminal organs to promote language development and learning.
[0074] In summary, the intonation-based auditory method effectively promotes English language learning. It provides sensory information to the brain through the body (vestibular system) and ears (auditory system), serving as the foundation for brain information processing and oral expression. Through vestibular training and vocalized body movements, it can actively improve learners' motor skills, spatial orientation, and memory span, achieving coordinated development of proprioception, hearing, and speech. In short, the intonation-based auditory method allows English learners to effectively control their spoken speech while listening to audio, aided by body movements. The listening materials used in this method consist of short English sentences covering eight sentence types: affirmative statements, negative statements, general questions, special questions, alternative questions, tag questions, imperative sentences, and exclamatory sentences. According to Gimson's Pronunciation of English, declarative sentences, special questions, imperative sentences, and exclamatory sentences generally end with a falling intonation; general questions end with a rising intonation; alternative questions use a rising intonation for the first choice and a falling intonation for the second; tag questions choose a falling intonation (indicating agreement between the speaker and listener) or a rising intonation (indicating no pressure on the listener to agree, expressing a simple inquiry) depending on the meaning. These eight sentence patterns cover different English intonation patterns, with 10 sentences in each pattern, totaling 80 sentences. All vocabulary in the sentences was selected from the 2000 words required by the 2022 edition of the Compulsory Education English Curriculum Standards. The audio materials were recorded by two native English speakers, one male and one female, using natural pronunciation and recorded in 32-bit stereo at a sampling rate of 44.1kHz using Adobe Audition CC (version 11.1.0), and all files are saved as *.wav files. Each short sentence was recorded for approximately 2000 milliseconds, containing about 5 syllables, at a speech rate of approximately 140 words per minute. Each sentence was read and recorded by two speakers, resulting in a total of 160 audio samples.
[0075] Low-pass filtering was also performed using Adobe Audition CC (version 11.1.0). A frequency band separator was used to separate the low-frequency speech from 0Hz to 300Hz in both channels. The low-frequency speech below 300Hz contains the fundamental frequency (F0), which determines the pitch variation in intonation. Based on the fundamental frequency curve of each sentence, this method sets the animation limits and motion trajectory of a 3D avatar to guide the user to make appropriate body movements in accordance with the intonation of the sentence.
[0076] AI-powered speech assessment technology involves speech input, feature extraction, and speech recognition based on acoustic and language models constructed using machine learning algorithms from a pre-existing speech and text database. After recognition, content analysis, pronunciation analysis, and prosody analysis are performed. Based on a manually labeled database, a machine assessment model trained using machine learning algorithms is used to score the speech, effectively reducing the cost of manual assessment.
[0077] The application of AI technology has solved many problems that human teachers cannot solve, thereby improving teaching efficiency.
[0078] Specifically, it includes:
[0079] a) AI is more accurate in processing low-frequency speech. Teachers rely on their own experience, and there will be differences in teaching depending on the teacher's level. AI, on the other hand, makes accurate responses based on audio data, which is both standardized and reduces human error.
[0080] b) AI teachers can solve the problem of timeliness, allowing learning anytime, anywhere.
[0081] c) AI teachers can conduct one-on-one training simultaneously, while human teachers cannot solve the problem of teaching multiple people simultaneously for oral English instruction.
[0082] d) The AI teacher can provide users with timely voice assessment results and feedback. The scoring criteria are based on the oral expression ability requirements of the *China Standards of English Language Ability* and the group standard *Computer Assessment Specifications for English Oral Proficiency Level Examination* (T / CIIA 009-2021). The system's scoring points include vocabulary pronunciation accuracy, stressed syllables, vocabulary stress, weak forms, linking, ellipsis, and intonation. The assessment results include phoneme scores, vocabulary scores, total sentence scores, total discourse scores, fluency scores, completeness scores, and prosody scores. The assessment is out of 100 points, with 80 points or above considered excellent, 60 points or above considered passing, and below 60 points considered failing. The AI assessment standards are unified, and it can identify the accuracy of phonemes, stress, and prosody. These require teachers to have a sufficiently high level of expertise to evaluate. Furthermore, the machine provides timely feedback, whereas teachers may need to listen multiple times to identify all the problems.
[0083] The AI teacher can provide each student with assessment results and a curve showing the change in their speech scores before and after. Based on each assessment result, especially the low scores, students can receive targeted training on key sentences / phonemes.
[0084] It should be noted that, for the interactive training method described in this invention, those skilled in the art will understand that all or part of the processes in the embodiments of this invention can be implemented by a computer program controlling the relevant hardware. The computer program can be stored in a computer-readable storage medium, such as in the memory of a server, and executed by at least one processor within the server. During execution, it can include the processes described in the embodiments of the information sharing method. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), etc.
[0085] Figure 1 and Figure 2 The spectrograms of the sentence "Can I have some soup?" are displayed separately for 300Hz low-pass filtered and unfiltered audio. Because the bandwidth of the low-pass filtered audio is limited, to ensure the same intensity, the amplitude of all low-pass filtered speech signals is normalized to 100%, and the amplitude of all unfiltered speech signals is normalized to 70%. Each sentence has two audio materials: one with low-pass filtering and one without, totaling 320 audio materials.
[0086] A data table is generated based on two columns of parameters. The x-axis represents time, and the y-axis represents frequency. The undefined data indicates that no frequency was detected at that time point. The other data are used to construct a scatter plot based on time and frequency, which can be seen to be a frequency plot with fluctuating characteristics.
[0087] like Figures 3-10 As shown
[0088] Based on the characteristics of declarative sentences, there is a rising intonation within the sentence and a falling intonation at the end. Taking the declarative sentence example "It's cold outside," where "it" specifically refers to the weather, the time interval corresponding to the stressed and rising intonation within the sentence is the waveform at x-coordinates 56-110 in the diagram; while the time interval corresponding to the falling intonation at the end of the sentence is the waveform at x-coordinates 118-152 in the diagram.
[0089] Based on the characteristics of general questions, the interrogative word has a rising intonation, and the sentence ending also has a rising intonation. In the general question above, "Can I have some soup?", the time interval corresponding to the interrogative word "Can I" is the waveform at x-coordinate 0-32 in the diagram. The word "soup" is stressed and pronounced with a rising intonation, which results in a rising intonation at the end of the sentence, corresponding to the time interval at x-coordinate 115-132 in the diagram.
[0090] Based on the characteristics of special interrogative sentences, the interrogative word and the key word have rising intonation, while the sentence ends with falling intonation. In the example of the special interrogative sentence above, "What time is it?", the time interval corresponding to the interrogative word "what" is the waveform at x-coordinate 0-16 in the diagram, the time interval corresponding to the stressed and rising intonation of the key word "time" is the waveform at x-coordinate 38-70 in the diagram, and the time interval corresponding to the falling intonation at the end of the sentence is the waveform at x-coordinate 71-118 in the diagram.
[0091] Based on the characteristics of imperative sentences, the imperative verb is stressed, and the entire sentence has a falling intonation. In the above imperative sentence, "Have some lunch, Mike," the stressed verb "have" corresponds to the time interval shown in the waveform graph at x-coordinate 0-18 in the diagram, while the falling intonation corresponds to the time interval shown in the waveform graph at x-coordinate 33-116 in the diagram.
[0092] Based on the characteristic of rising and falling tones in each sentence type and the fundamental frequency value extracted from the audio, the system is designed with 3D animation to present the rising and falling tones of each speech sentence in order to better help students train their intonation through body movements. The 3D animation provides guidance to students and helps them imitate, thus enabling them to better complete this training method.
[0093] Based on the characteristics of negative declarative sentences, the negative word and key words in the sentence are stressed and pronounced with a rising intonation. Taking the negative declarative sentence "You didn't come to school" as an example, the negative word "didn't," as the meaning to be emphasized in the sentence, is stressed and pronounced with a rising intonation, corresponding to the time interval shown in the waveform graph at x-coordinates 175-419 in the diagram. The key word "school" is also stressed and pronounced with a rising intonation, corresponding to the time interval shown in the waveform graph at x-coordinates 1079-1387 in the diagram.
[0094] Based on the characteristics of tag questions, the subject and keyword exhibit emphasis and rising intonation. Taking the tag question "Jane has never been to America, has she?" as an example, the subject "Jane" is stressed and has a rising intonation, corresponding to the time interval shown in the waveform graph at x-coordinate 0-0.287. The keyword "America" is also stressed and has a rising intonation, corresponding to the time interval shown in the waveform graph at x-coordinate 1.06-2. The tag question at the end of the sentence uses a rising intonation to inquire about the other person's opinion and confirm one's own judgment.
[0095] Based on the characteristics of alternative questions, both the interrogative word and the key word exhibit an upward intonation. Taking the alternative question "Can you sing or dance?" as an example, the interrogative word "Can you" has an upward intonation, corresponding to the time interval shown in the waveform diagram at x-coordinate 5-509. The interrogative word "sing" also has an upward intonation, corresponding to the time interval shown in the waveform diagram at x-coordinate 707-960. The alternative interrogative word "or" has a downward intonation and is pronounced weakly, corresponding to the time interval shown in the waveform diagram at x-coordinate 135-165. The other part of the alternative question, "dance," is emphasized again with an upward intonation, corresponding to the time interval shown in the waveform diagram at x-coordinate 183-236.
[0096] Based on the characteristics of exclamatory sentences, interjections and keywords are stressed and have an upward intonation. Taking the exclamatory sentence example "What an interesting film!" above, the interjection "what" is stressed and has an upward intonation, corresponding to the time interval shown in the waveform graph at x-coordinate 0-454 in the diagram, and the keyword "interesting film" is stressed and has an upward intonation, corresponding to the time interval shown in the waveform graph at x-coordinate 744-1773 in the diagram.
[0097] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. A method for training auditory speech in 3D virtual avatars based on AI interaction, characterized by: Includes the following steps: S1. Record the speech signal and perform low-pass filtering on the original speech to retain low-frequency audio below 300Hz. The low-frequency audio retains the rhythmic features of the speech. S2. The processed input speech signal is divided into eight speech sentence types: affirmative statement, negative statement, general question, special question, alternative question, tag question, imperative sentence, and exclamatory sentence. The user is continuously trained for no less than 30 seconds. The audio of the sentences used for training is processed with low-pass filtering to highlight speech parameters, which can enhance the user's ontological perception of the language signal. In this training, through sensory integration training of senses, hearing, and vision, the learner can maximize the development of brain neural pathways and expand their learning potential. S3. Create a virtual image. Create a 3D virtual image in Unity 3D. Match the skeletal structure of the 3D virtual image with the predefined skeletal structure of Mecanim. Create an Animator component for each action to build the animation effect of waving arms. Each action is triggered by the controller instruction. For each action effect, the corresponding parameters need to be set to control the amplitude of the action. S4. Integrate the sentence type with the corresponding 3D virtual character. The 3D virtual character has two forms of movement: one is the up and down movement of the arms, and the other is the opening and closing of the hands. By obtaining the highest and lowest points of the audio fundamental frequency, set the highest point of the arm waving or the maximum degree of the hands opening in the 3D virtual character as the highest point of the audio fundamental frequency; set the lowest point of the arm waving or the minimum degree of the hands closing in the 3D virtual character as the lowest point of the audio fundamental frequency. Based on the current audio frequency point, the amplitude of the 3D virtual character is 100% * (current frequency – minimum frequency) / (maximum frequency – minimum frequency) = amplitude%. During the actual animation playback, users can choose between two animation formats, or the system can randomly display one. The system obtains the current time in real time, and the 3D virtual character displays the animation based on the time and the calculated movement amplitude. When an undefined value is encountered, the animation does not change until a valid frequency parameter is obtained. S5. Playing the speech signal in units, the low-frequency speech of the two-channel 0Hz to 300Hz is separated by a frequency band separator. This includes the fundamental frequency that determines the pitch. The fundamental frequency determines the pitch change of the intonation. Based on the fundamental frequency curve of each sentence, this method sets the animation limit value and motion trajectory of the 3D virtual image to guide the user to make appropriate body rhythms with the intonation of the sentence. S6. The words in the sentence type are divided into high and low tone words and displayed as word levels according to their pitch. While listening to the low-pitched filtered audio material, the user watches the 3D virtual image. In this step, the user needs to imitate the low-pitched filtered voice and make body movements in reference to the 3D virtual image. The 3D virtual image will use the rhythm of the sentence as a melody and make body movements in accordance with the stress, rhythm and intonation of the sentence. In this step, the user imitates the low-pitched filtered voice while making body movements. S7. AI Assessment: The AI assessment module assesses the user's pronunciation. The assessment module includes a display module for intonation, actions, and their combinations; a module for marking intonation, actions, and their combinations to form an assessment; a module for scoring the marked results; a result output module; and a module for assessment results and suggestions.
2. The AI-interactive 3D virtual character auditory speech training method according to claim 1, characterized in that: It includes a data acquisition module for acquiring voice data and motion data from multiple users; the voice data is collected by a voice acquisition device and includes a 3D virtual image seen from the user's perspective; the motion data is collected by a motion acquisition device and includes vector data generated by the user's motion. The scenario generation module is used to select and generate training scenarios from multiple preset training scenarios in response to user actions. The role matching module is used to determine the content that the user sees in the training scenario based on the user's voice data, and to match the user's corresponding hand movements from the training scenario based on the content seen. The motion generation module is used to generate the motion corresponding to the 3D virtual image in the training scenario based on the user's motion data. The panoramic integration module is used to integrate multiple users and actions and output them in panoramic mode; The output in panoramic mode includes: when performing panoramic integration, determining the aspect ratio of the panoramic mode based on the degree of user dispersion; or when performing panoramic integration, extracting videos containing users from the scene and stitching the videos containing users together. The feedback module is used to determine whether the difference between the user's action and the reference action is within a preset range, and when it is not within the preset range, it generates feedback information; the feedback information includes reminding the user of the current improper operation through vibration, specifically: using a patch on the part corresponding to the action that is not within the preset range to remind the user of improper movement.
3. The AI-interactive 3D virtual character auditory speech training method according to claim 2, characterized in that, The data acquisition module includes: The identification unit is used to generate a corresponding user identification number for each user. The matching unit is used to match the user identifier with the corresponding eye-tracking acquisition device and motion acquisition device; The acquisition unit is used to generate eye movement data and motion data for each user through the eye movement acquisition device and the motion acquisition device.
4. The AI-interactive 3D virtual character auditory speech training method according to claim 3, characterized in that, The data acquisition module is further configured to acquire audio data from multiple users; the panoramic integration module is further configured to integrate the audio data with the user's corresponding actions, and also includes: An audio playback module is used to receive hand gestures from the 3D virtual avatar or the user and play the corresponding audio data.
5. The AI-interactive 3D virtual character auditory speech training method according to claim 4, characterized in that, The data acquisition module is also used to acquire audio data from multiple users; The audio adjustment module is used to integrate the actions and audio data of the multiple users, adjust the volume according to the distance of the user's actions in the panoramic mode, and store the data.
6. The AI-interactive 3D virtual character auditory speech training method according to claim 5, wherein the stored data is detected and scored through AI evaluation.
Citation Information
Patent Citations
China sign language standardization train learning method based on video and three-dimensional route planning
CN102663927A
Spoken dialogue system, a spoken dialogue method and a method of adapting a spoken dialogue system
US20180226076A1