A lip reading auxiliary training system based on a meta universe and application thereof
The Metaverse-based lip-reading learning assistance training system, utilizing lip-reading training modules and virtual human dialogue communication modules, solves the problems of low accuracy and insufficient interactivity in existing lip-reading learning technologies, thereby improving the communication ability and learning outcomes of hearing-impaired individuals.
Patent Information
- Application Number
- CN202310371018.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-07
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2043-04-07
AI Technical Summary
Existing lip-reading learning technologies suffer from low accuracy, untimely feedback, lack of user interaction, and difficulty for users to enhance their real-world communication skills through system assistance.
This paper presents a metaverse-based lip-reading learning assistance training system, which includes a lip-reading training module, a virtual human dialogue and communication module, and a user personal center module. The system conducts lip-reading training through metaverse learning scenarios, uses virtual humans for dialogue and communication, and provides personalized feedback and interaction.
It improves the accuracy of lip reading, enhances users' communication skills, provides a variety of learning methods and scenarios, increases users' learning motivation and interactivity, and meets the psychological needs of hearing-impaired individuals.
Smart Images

Figure CN116524791B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of lip-reading learning, and more particularly relates to a lip-reading learning auxiliary training system based on a meta-universe and application thereof. BACKGROUND
[0002] Hearing-impaired people are an indispensable part of society. In the past, a lot of work has been done to help hearing-impaired people better integrate into social life. However, these achievements either require complex and expensive equipment or are difficult to use and not widely applicable. In addition, previous achievements generally lack care for the minds of hearing-impaired people and are difficult to effectively solve the problem of assisting deaf-mutes to integrate into society. Therefore, the construction of a hearing-impaired auxiliary system that is simple in equipment, convenient to operate, easy to master, and provides humanistic care plays a very important role in improving the quality of life, social participation, and happiness of hearing-impaired people.
[0003] The existing auxiliary hardware system includes:
[0004] (1) Sign language translation gloves
[0005] a. Through five sensors for collecting joint motion states at the joint of the gloves, simple information output is realized after processing by a flexible circuit module, specific information can be collected through gesture motion, and the signal is processed and recognized by the flexible circuit module, and the recognized information is sent out through a wireless communication module, received by a terminal device and displayed in the form of pictures or audio.
[0006] b. Advantages: The device is relatively light and has less impact on normal work and life.
[0007] c. Disadvantages: The cost of learning sign language, compared to the use of gestures by able-bodied people as an aid to better express information in communication, the use of sign language limits the user's gesture expression in the communication process. In addition, due to the low popularity of sign language among the general public, the standard of sign language is not unified, which limits the communication objects.
[0008] (2) Speech recognition device
[0009] a. Patsnap, Microsoft, Baidu, etc.
[0010] b. Advantages: It can conveniently realize the expression of normal people to deaf-mutes.
[0011] c. Disadvantages: It can only improve the efficiency of normal people expressing to deaf-mutes, and it is difficult to solve the problem of communication difficulties between deaf-mutes and normal people. Deaf-mutes still need to use typing, sign language, etc. to express themselves, which limits the communication efficiency.
[0012] (3) Cochlear implant
[0013] a. Cochlear, Sonova, Nurotron.
[0014] b. Advantages: The sound is converted into a certain coded form of electrical signal by the external speech processor, and the auditory function of the deaf is restored or reconstructed by directly exciting the auditory nerve through the implanted electrode system, which can communicate with normal people.
[0015] c. Disadvantages: The equipment of cochlear implant needs regular maintenance and cleaning, and the service life is limited. It needs to be installed through surgery, which is risky and expensive. In addition, implanting cochlear implant can cause a series of complications, such as subcutaneous hematoma, acute otitis media, etc., which brings more pain to patients.
[0016] (4) Hearing aid
[0017] a. iCom, Sivantos.
[0018] b. Advantages: Small loudspeaker, amplify the sound that can not be heard, use the residual hearing of hearing-impaired people to send the sound to the auditory center of the brain, so as to feel the sound.
[0019] c. Disadvantages: A device to improve hearing, if the user loses hearing completely, the product is invalid, and there are limitations on the use of groups; secondly, while enlarging the effective sound, noise will also be enlarged, users will hear a lot of noise, and cannot guarantee the use effect in all scenarios.
[0020] According to the investigation, the existing systems related to lip reading on the market are mainly divided into two types: pure lip reading system and lip teaching auxiliary system.
[0021] The pure lip reading system mainly aims at public security, disabled education, and identity recognition, and pushes out a system taking lip reading technology as the core. At present, there are mainly the lip reading system pushed out by Sogou, which captures the lip shape of the user through the app combined with the mobile phone camera, and expects to transplant the technology to more fields in the future; the intelligent department of mechanical engineering of Tsinghua University jointly with the biomechanical team pushes out a novel lip reading system, which can accurately recognize the lip reading through recognizing the human facial muscle movement after excluding the angle, light, and shielding and other external environmental factors caused by the camera; and Haiyun Data also pushes out a set of lip reading system concept combining big data visualization analysis and AI technology to enter the silent data recognition task of the public security department. The pure lip reading system mainly pursues the lip information, and achieves the purpose of converting the lip reading into text or information that the receiver can understand. Such system mainly focuses on the accuracy of the lip reading technology, and does not constitute a complete human-computer interaction ecological system. Such system lacks the interactivity of human-computer interaction, and at present, lacks the application scene for training the hearing-impaired people, and cannot provide the training and auxiliary teaching required by the hearing-impaired people, but only can convert the lip shape and text, and cannot fundamentally solve the problems of the hearing-impaired people.
[0022] The lip teaching auxiliary system mainly aims at the application scene of disabled education, and pushes out a teaching auxiliary system taking the standard database as the core. At present, there are companies in China that have proposed a "three-dimensional lip interactive teaching system". The system includes a knowledge base including text, vocabulary, and lip shape. The standard lip shape and necessary semantic knowledge are displayed and demonstrated to students through three-dimensional animation. The disabled students will be able to assist their own lip learning by comparing with the standard knowledge base. The lip teaching system currently pushed out mainly provides the standard lip shape library, but lacks the human-computer interaction, and the students cannot get feedback through the system, so as to know whether their lip shape is correct or not, and the auxiliary effect on learning is not strong. Due to the lack of interaction, the teaching mode of the system is rigid, and it is no different from reading from a book. The emotional auxiliary needs of the hearing-impaired people cannot be met.
[0023] In summary, the existing lip learning technology has the technical problems of low accuracy, untimely feedback, lack of interaction with the user, and difficulty for the user to enhance the communication ability in reality through the system. SUMMARY
[0024] In view of the above defects or improvement needs of the prior art, the present application provides a lip learning auxiliary training system based on a meta universe and an application thereof, thereby solving the technical problems of the existing lip learning technology, such as low accuracy, untimely feedback, lack of interaction with the user, and difficulty for the user to enhance the communication ability in reality through the system.
[0025] To achieve the above object, according to one aspect of the present application, a meta-universe-based lip language learning auxiliary training system is provided, comprising a lip reading training module, a virtual person answering communication module and a user personal center module.
[0026] The lip reading training module is configured to store pre-acquired standard lip shape videos, establish a meta-universe learning scene, and enable a user to perform lip reading training through the standard lip shape videos in the meta-universe learning scene. The lip reading training module is further configured to identify the text of the user's lip reading from the lip language learning video of the user's lip reading training through the standard lip shape videos, calculate the similarity between the text of the user's lip reading and the text of the standard lip shape videos, and determine the lip reading training effect of the user through the similarity.
[0027] The virtual person answering communication module is configured to establish a meta-universe social scene, identify social text from the video of the user's speech in the meta-universe social scene, and convert the answering text of the social text in the answering process into audio to form a virtual person in combination with a face, so that the user can communicate with the virtual person in the meta-universe social scene.
[0028] The user personal center module is configured to record and feedback the lip reading training effect of the user, combine the audio of the user with the face to form a virtual image of the user, and enable the user to communicate with other users using the lip language learning auxiliary training system in the meta-universe social scene in the form of the virtual image.
[0029] Further, the lip reading training module comprises a video preprocessing module, a lip language recognition module and a feedback module,
[0030] The video preprocessing module is configured to store pre-acquired standard lip shape videos in multiple languages, and clip the standard lip shape videos in each language into standard lip shape videos in word mode and sentence mode.
[0031] The lip language recognition module is configured to identify the text of the user's lip reading from the lip language learning video of the user's lip reading training through the standard lip shape videos in word mode or sentence mode in different languages.
[0032] The feedback module is configured to calculate the similarity between the text of the user's lip reading and the text of the standard lip shape videos, and feed back to the user personal center module.
[0033] Further, the lip reading training module further comprises a lip language recognition model,
[0034] The lip language recognition model comprises a front-end feature extraction network and a back-end classification network, and is trained in the following manner:
[0035] The video frame is a video frame of different languages, and finally a lip reading model of different languages is obtained.
[0036] The video frame is a video frame of different languages, and finally a lip reading model of different languages is obtained.
[0037] The lip reading module is configured to recognize the text of the user's lip reading from a lip reading training video of the user in a word mode or a sentence mode of a standard lip shape video in a language using a lip reading model of the language.
[0038] Further, the virtual human answering communication module comprises a virtual human forming module and a dialogue robot,
[0039] The virtual human forming module is configured to call the lip reading model to recognize social text from a video in which the user speaks in a meta-universe social scene, input the social text into the dialogue robot, and convert the answer text output by the dialogue robot into audio to form a virtual human in combination with a face.
[0040] Further, the virtual human forming module comprises a speech synthesis module and an animation generation module,
[0041] The speech synthesis module is configured to synthesize audio from the text output by the dialogue robot through a speech synthesis software.
[0042] The animation generation module is configured to combine the audio with the face using a speaker face generation model to form a virtual human; wherein the speaker face generation model comprises an encoder, a decoder and a lip shape discriminator, and the speaker face generation model is trained in the following manner:
[0043] The sample speech segment is converted into a mel spectrum form, the sample speech segment in the mel spectrum form is encoded into preprocessed audio through residual convolution in the encoder, the sample face picture is down-sampled through residual convolution in the encoder to obtain a preprocessed face picture, the preprocessed audio and the preprocessed face picture are decoded through transpose convolution in the decoder to form a virtual human; the lip shape discriminator encodes the lip shape and the audio of the virtual human through two convolution networks, and the training is performed to convergence with the minimum error between the encoded lip shape and the lip shape in the preprocessed face picture and the minimum error between the encoded audio and the preprocessed audio as the target, to obtain the trained speaker face generation model.
[0044] Further, the dialogue robot is a personalized dialogue robot, and the dialogue robot is personalized by the following manner:
[0045] The dialogue text of a psychological counselor or a teacher of a school for the hearing impaired is collected, and before the user dialogues with the dialogue robot, the dialogue text is input into ChatGPT, Wenxin Yiyang, WeChat bug hole assistant, chat robot PET, chat robot Bard or chat robot MOSS to guide the dialogue robot to play the role of a psychological counselor or a teacher of a school for the hearing impaired.
[0046] Further, the lip learning auxiliary training system further comprises a metaverse scene establishing module,
[0047] The metaverse scene establishing module is configured to establish a metaverse scene using Multispace or MetaStack of Baidu Xilang metaverse base;
[0048] The lip reading training module is configured to call the metaverse scene establishing module to establish a metaverse learning scene;
[0049] The virtual human dialogue communication module is configured to call the metaverse scene establishing module to establish different metaverse social scenes;
[0050] The virtual human forming module is configured to identify social text in a video in which a user speaks in different metaverse social scenes, input the social text into a dialogue robot, and combine audio output by the dialogue robot with a face to form a virtual human in different metaverse social scenes, so that the user dialogues with the virtual human in the corresponding metaverse social scene in the different metaverse social scenes.
[0051] Further, the user personal center module is configured to store and manage video data of lip learning of a user using the lip learning auxiliary training system, call the virtual human forming module to combine audio of the user with a face to form a virtual image of the user, and call the metaverse social scene establishing module to establish a metaverse private space of the user, so that the user communicates with other users using the lip learning auxiliary training system in the metaverse private space.
[0052] According to another aspect of the present application, an application of a meta-universe-based lip-reading learning auxiliary training system is provided, which is applied to assist hearing-impaired persons in learning lip-reading. The hearing-impaired persons, as users of the lip-reading learning auxiliary training system, select standard lip shape videos from a lip-reading training module to perform lip-reading training in a meta-universe learning scene, and the similarity output by the lip-reading training module is used to judge the lip-reading training effect of the users. The users select virtual persons from a virtual person answering communication module, and perform answering communication with the virtual persons in a meta-universe social scene. The users select a user personal center module to customize a virtual image, and perform answering communication with other users using the lip-reading learning auxiliary training system in the meta-universe social scene.
[0053] According to another aspect of the present application, an electronic device is provided, characterized by comprising:
[0054] a memory having a computer program stored thereon;
[0055] a processor configured to execute the computer program in the memory to implement processing steps of a meta-universe-based lip-reading learning auxiliary training system.
[0056] Overall, compared with the prior art, the above technical solutions conceived by the present application can achieve the following beneficial effects:
[0057] (1) The present application first applies virtual scene and virtual person technology to a lip-reading auxiliary system, increasing the system interaction. In the present application, the lip-reading training module provides a meta-universe learning scene and standard lip shape videos for the user to perform lip-reading training, the virtual person answering communication module provides a meta-universe social scene for the user, so that the user can perform answering communication with the virtual person in the meta-universe social scene, and the user personal center module allows the user to perform answering communication with other users using the lip-reading learning auxiliary training system in the meta-universe social scene. The present application system provides multiple ways and scenes for the user to learn lip-reading, and the learning auxiliary method is diverse, thereby improving the accuracy of lip-reading learning and enhancing the user's communication ability in the real world through system assistance. The similarity can be used to timely feedback the lip-reading training effect of the user. In the virtual person answering communication module, the user can enter the meta-universe social scene and immerse in the atmosphere to trigger a dialogue with the virtual person. On the one hand, this can arouse the interest of the user and increase the time of using lip-reading, which is helpful for the user to further master lip-reading. On the other hand, the user can immerse in the virtual space and try to communicate without burden, and the positive social activities at any time and any place are conducive to increasing the motivation of the user to practice and use lip-reading. In the user personal center module, the user can customize a personal virtual image, understand the learning effect, and increase the interaction with other users.
[0058] (2) The application provides videos in different languages and different learning modes for users to learn lip language through a video preprocessing module, provides a variety of learning content, and expands the audience group of the lip language learning auxiliary training system. The training data of the lip language recognition model in the application is video frames in different languages, and the lip language recognition model in different languages is obtained, so that the accuracy can be improved when lip language recognition in different languages is performed. In the training, the ROI sequence and the differential ROI sequence are respectively input into two branches of the front-end feature extraction network, the original input part is retained, a new branch is used to extract features from the differential data, and finally the two are added to fuse the information of the two, so that the motion feature capture ability of the model is enhanced while the features of each frame are extracted. The application can accurately recognize lip language in a non-restricted environment. At the same time, the model recognition accuracy is high, the recognition efficiency is high, and the generalization performance is good.
[0059] (3) The virtual person forming module of the application identifies social text from the video of the user speaking in the metaverse social scene, inputs the social text into the dialogue robot, and combines the audio after converting the answer text output by the dialogue robot into audio with the face to form a virtual person. A virtual person is formed in the metaverse social scene, which can respond to the scene and better communicate with the user, improving the user experience. When synthesizing the virtual person, the mouth shape discriminator reduces the mouth shape and audio error, improves the mouth shape effect, and solves the problem of unsatisfactory mouth shape effect of the past model.
[0060] (4) The dialogue robot in the application can be personalized by adjusting a variety of existing robots, guiding the dialogue robot to simulate a psychological consultant or a teacher in a school for the hearing impaired, and customizing for the needs of the hearing impaired to better meet the psychological needs of the hearing impaired. While meeting the daily communication needs, it provides psychological comfort and support for the hearing impaired, reduces their stress, restores their confidence, and protects the mental health of the hearing impaired.
[0061] (5) The meta universe scene can be established in various ways, and the lip reading training module, the virtual person answering communication module and the user personal center module can all call the meta universe scene establishment module to establish the required virtual scene. The virtual person answering communication module calls the meta universe scene establishment module to establish different meta universe social scenes, and the virtual person forming module forms different virtual persons for different meta universe social scenes. The user can select any scene and immerse in it, and dialog with each virtual person in the scene, thereby virtually practicing lip language. The meta universe social scene provides a new communication method and experience for the user, so that the user can immerse in the virtual space and try to open the mouth to communicate more freely, and more naturally carry out social activities. The meta universe social activities that can be carried out anytime and anywhere are also more conducive to promoting the lip shape correction and lip language practice of the hearing impaired group, and improving the lip language training motivation of the hearing impaired group, so as to form a virtuous cycle.
[0062] (6) The user personal center module in the application can store and manage data, view user practice duration and view user practice effect. The module can help the user better understand the learning progress and learning situation. Custom personal virtual image, establish meta universe private space, build social scene belonging to the user, so that the user communicates with other users using the lip language learning auxiliary training system in the meta universe private space, and creates a new social scene under the meta universe.
[0063] (7) The lip language learning auxiliary training system designed in the application is applied to assist the hearing impaired to learn lip language. In the lip language training module, the user can watch the lip shape of the standard lip shape video to imitate the lip shape and practice lip language pronunciation. The user can continuously practice to improve the lip reading accuracy by comparing the lip movement with the standard lip movement. In the virtual person answering communication module, the system builds a communication platform based on the virtual person, and the user can dialog with the virtual person. The user can immerse in the virtual space and try to open the mouth to communicate more freely, and the positive social activities anytime and anywhere are conducive to increasing the user's motivation to practice and use lip language. In the user personal center module, the user can individualize the virtual person image according to the user's image, and view the user practice effect. The user can enhance the communication ability in the real world through the system assistance. BRIEF DESCRIPTION OF DRAWINGS
[0064] Figure 1 is the architecture diagram of the entire system and the internal modules of the system provided by the embodiment of the application;
[0065] Figure 2 is the internal logic flowchart of the lip reading training module provided by the embodiment of the application;
[0066] Figure 3 is the internal logic flowchart of the virtual person answering communication module provided by the embodiment of the application;
[0067] Figure 4 is a logical flowchart of virtual human technology implementation provided by the embodiment of the present application;
[0068] Figure 5 is an internal logical flowchart of the user personal center module provided by the embodiment of the present application. DETAILED DESCRIPTION
[0069] In order to make the purpose, technical solutions and advantages of the present application clearer and more apparent, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0070] As shown in Figure 1 , a lip-reading learning auxiliary training system based on the meta universe, characterized in that it comprises a lip-reading training module, a virtual human answering communication module and a user personal center module.
[0071] The lip-reading training module is used to store pre-acquired standard lip shape videos, establish a meta universe learning scene (classroom, study, library or office), and enable the user to perform lip-reading training through the standard lip shape videos in the meta universe learning scene. The text of the user's lip-reading is recognized from the lip-reading learning video of the user through the standard lip shape videos, the similarity between the text of the user's lip-reading and the text of the standard lip shape videos is calculated, and the lip-reading training effect of the user is determined by the similarity.
[0072] The virtual human answering communication module is used to establish a meta universe social scene, recognize social text from the video of the user speaking in the meta universe social scene, convert the answer text in the answering process of the social text into audio and combine it with the face to form a virtual human, so that the user can communicate with the virtual human in the meta universe social scene.
[0073] The user personal center module is used to record and feedback the lip-reading training effect of the user, combine the audio of the user with the face to form a virtual image of the user, and enable the user to communicate with other users using the lip-reading learning auxiliary training system in the meta universe social scene in the form of a virtual image.
[0074] Embodiment 1
[0075] The use of the lip-reading learning auxiliary training system by the user is described in detail in embodiment 1.
[0076] After the user enters the system, the user can select any one of the lip-reading training module, the virtual human answering communication module and the user personal center module.
[0077] The user enters the lip-reading training module, the standard lip shape video pre-acquired in the lip-reading training module is saved in the database, and the lip language learning video recorded by the user through the standard lip shape video for lip-reading training in the meta-universe learning scene is also saved in the database after the user agrees. The system of the application can be a VR glasses with a camera on the hardware, which can display a 3D panorama and record the facial expressions of the user, capture the facial movements of the user, and thus record the user's own video to obtain the lip language learning video of the user. The text of the user's lip-reading is recognized from the lip language learning video, and the similarity between the text of the user's lip-reading and the text of the standard lip shape video is calculated. When the similarity is less than a preset value, the user continuously performs lip-reading training through the standard lip shape video in the meta-universe learning scene to provide the lip-reading accuracy.
[0078] The user enters the virtual human answering communication module, and the user can perform answering communication with the virtual human in the meta-universe social scene, or perform answering communication with other users in the meta-universe social scene.
[0079] The user enters the user personal center module, can view the lip-reading training effect of the user, customize the virtual image of the user, and make the user perform answering communication with other users using the lip language learning auxiliary training system in the meta-universe social scene in the virtual image.
[0080] Embodiment 2
[0081] The answering communication in the case that the user is trained to be qualified is described in detail through Embodiment 2.
[0082] The user enters the lip-reading training module, and the user records a lip language learning video by performing lip-reading training through the standard lip shape video in the meta-universe learning scene. The text of the user's lip-reading is recognized from the lip language learning video, and the similarity between the text of the user's lip-reading and the text of the standard lip shape video is calculated.
[0083] When the similarity output by the lip-reading training module is less than a preset value, the user performs lip-reading training by acquiring the standard lip shape video from the lip-reading training module, and when the similarity output by the lip-reading training module is greater than or equal to the preset value, the user performs answering communication with the virtual human or other users in the meta-universe social scene in the real image or the virtual image.
[0084] Embodiment 3
[0085] The functions of the lip-reading training module and the case that the user uses the lip-reading training module are described in detail through Embodiment 3.
[0086] The lip-reading training module comprises a video preprocessing module, a lip language recognition module and a feedback module,
[0087] The video preprocessing module is configured to store standard lip shape videos in multiple languages collected in advance, and clip the standard lip shape videos in each language into standard lip shape videos in word mode and sentence mode.
[0088] The lip reading training module is configured to recognize the text of the user's lip reading from the lip language learning video of the user's lip reading training through the standard lip shape videos in word mode or sentence mode in different languages.
[0089] The feedback module is configured to calculate the similarity between the text of the user's lip reading and the text of the standard lip shape video, and feed back to the user personal center module.
[0090] The lip reading training module further comprises a lip language recognition model,
[0091] The lip language recognition model comprises a front-end feature extraction network and a back-end classification network, and is obtained by training in the following manner:
[0092] The face image and the real lip language in the video frame are obtained, the lip region of the face image is extracted, and a ROI sequence is formed. The ROI sequence and the differential ROI sequence are input into two branches of the front-end feature extraction network respectively, and the lip region feature of the spliced differential feature is output. The lip region feature of the spliced differential feature is input into the back-end classification network, and the predicted character is output. The error between the predicted character and the real lip language is minimized as the target, and the training is performed until convergence to obtain the lip language recognition model.
[0093] The video frame is a video frame in different languages, and finally a lip language recognition model in different languages is obtained.
[0094] The lip language recognition module is configured to recognize the text of the user's lip reading from the lip language learning video of the user's lip reading training through the standard lip shape videos in word mode or sentence mode in different languages.
[0095] As Figure 2As shown, when the user selects to enter the lip-reading training module, language selection is first performed, and if Chinese is selected, standard Chinese lip shape database (composed of standard Chinese lip shape videos) is used for training, and if English is selected, standard English lip shape database (composed of standard English lip shape videos) is used for training, and then training mode is selected, word mode or sentence mode, after the user selects the training mode suitable for himself / herself, the standard Chinese lip shape database or the standard English lip shape database provides learning videos, and the user can clearly see the 3D model of the standard lip shape, and can more accurately imitate and learn. At the same time, through the user's own video, after learning, the user personal center module is uploaded, which is compared with the standard video to obtain the similarity, so that the user can improve the lip-reading ability in the interactive feedback. In order to serve the Chinese lip shape training, the lip-reading model is designed, Chinese data can be used in training to improve the accuracy of Chinese lip-reading, and the feedback mechanism of the model is used to provide the lip shape correction function.
[0096] The source of the standard Chinese lip shape video in the standard Chinese lip shape database has three parts: 1. standard Chinese news broadcast video, 2. Chinese recording video of teachers in a lip school, and 3. common life scene video, which can be derived from standard Mandarin movies and TV series. The standard Chinese lip shape video can be edited into word mode and sentence mode.
[0097] The source of the standard English lip shape video in the standard English lip shape database has three parts: 1. standard English news broadcast video, 2. English recording video of teachers in a lip school, and 3. common life scene video, which can be derived from English movies and English TV series. The standard English lip shape video can be edited into word mode and sentence mode.
[0098] In a similar way, standard Japanese, Korean, German or French lip shape databases can also be obtained.
[0099] The user can continuously practice and improve the lip-reading accuracy in multiple videos by comparing his / her own lip movement with the standard lip movement.
[0100] Embodiment 4
[0101] The function of the virtual human Q&A communication module and the use of the virtual human Q&A communication module by the user are described in detail through Embodiment 4.
[0102] As Figure 3As shown, the virtual human interactive communication module provides a truly simulated reality communication platform for the hearing impaired. On the one hand, it solves the problem of lack of lip-reading teachers in real learning process and lack of practice communication objects. On the other hand, due to the long-term lack of effective communication with the outside world, the hearing impaired often form a closed circle centered on themselves and are out of touch with society. The virtual human interactive communication module can have effective communication with psychological and emotional comfort, giving the hearing impaired an opportunity to open up.
[0103] When the user selects the virtual human interactive communication module, the system will provide different meta-universe social scenes, and the user can immerse in the real world with emotional communication. In order to realize the function of the module, the present application builds a communication platform based on virtual human and ChatGPT technology (here it can also be Wenxin Yiyang, WeChat bug hole assistant, chat robot PET, chat robot Bard or chat robot MOSS).
[0104] As shown in Figure 4 The text result obtained by the lip-reading model is input into the QA module (question and answer module) and the TTSA (text-to-speech and animation) module to generate a virtual human. The QA module is composed of a fine-tuned ChatGPT (here it can also be Wenxin Yiyang, WeChat bug hole assistant, chat robot PET, chat robot Bard or chat robot MOSS), which answers in real time through streaming technology. The TTSA module is composed of two parts: speech synthesis and speaker face generation. The speech synthesis uses Microsoft Azure speech synthesis API to convert the text content generated by the QA module into audio for animation generation. The speaker face generation is based on the wav21ip model and optimizes the lip movement to generate vivid and accurate speaker animation.
[0105] First, the lip-reading algorithm is used to recognize the user's mouth shape, recognize the content of the speech and output it as a text format. The text is input into the QA module. The QA module is composed of a ChatGPT fine-tuned for the needs and characteristics of the hearing impaired. ChatGPT is a dialogue robot based on GPT-3 further trained by OpenAI. It has rich dialogue content and the ability to have continuous dialogue. The original ChatGPT system can realize simple question and answer, daily chat and other functions. Through appropriate instructions, ChatGPT can simulate a psychological consultant, a hearing impaired school teacher, a psychological care worker, a virtual companion assistant, etc. to meet the psychological needs of the hearing impaired.
[0106] Using appropriate prompt sentences, the ChatGPT is prompted to understand the emotions and psychology of the hearing-impaired person and to make polite and emotional answers when communicating with the hearing-impaired person.
[0107] The dialogue robot is trained in the following way:
[0108] The first step is to collect a series of questions and manually answer them, and fine-tune the GPT-3 model with these questions and answers; the second step is to let the fine-tuned model answer the above questions, and generate multiple answers for each question, and manually sort these answers from high to low according to quality, which is used to train the reward model (reinforcement learning term); the third step is to make answers by the fine-tuned GPT-3, and the reward model generates the results according to the answers, and further optimizes through reinforcement learning.
[0109] When calling the API of ChatGPT, it usually starts with an initial command, which contains system, user and assistant. The system command indicates the specific identity of the assistant, the reply tone, the function, etc.: such as "you are a psychological counselor, you should listen carefully to the user's speech, understand the user's emotions, empathize with the user, make gentle and peaceful, considerate answers, and provide some suggestions for the user's troubles, and comfort the user's heart." Then, provide one or more examples of user and assistant dialogue to further clarify the role to be played. Finally, the user's input is transmitted into the system and the dialogue with ChatGPT begins. At this time, ChatGPT has a full understanding of the role to be played and can make satisfactory answers.
[0110] Some sensitive content is shielded by black list. Trigger exception result directly. Judge the emotion of the answer, when the answer is negative, call the API again to generate a new answer, until the answer emotion meets the requirements, if multiple calls fail, return an exception result and submit the exception log to the maintenance personnel for investigation.
[0111] The application uses Multispace multi-space or Baidu Xilang metaverse base MetaStack to establish a metaverse scene.
[0112] Specifically, Multispace provides users with a variety of construction tools, trying to upgrade from UGC to AIGC. Currently, the platform provides visual UGC editing tools that can be dragged and dropped to design and build metaverse buildings and characters, etc.; for more professional users, it provides SDKs that can implement more interactive functions; for users who pursue simplicity, it provides a building transaction market that can be ordered and deployed quickly; users can also sell their own designed buildings, applications, and artworks on the transaction market. In addition, the platform is also trying to provide AIGC tools for users, allowing future product interactions to move from text and image clicks to voice commands, making it easier to create metaverse scenes.
[0113] Baidu released the MetaStack metaverse base, which is based on a series of metaverse infrastructure and a one-stop development platform. It can create an independent metaverse in as little as 40 days, greatly reducing the time cost of metaverse development. As the first domestic metaverse platform, it not only has basic conference systems, exhibition art centers, digital collectibles, and metaverse auctions, but also has powerful "AI + cloud computing" to better handle massive data processing and super-large model training. And to solve the problems of low R&D efficiency and high operating costs, it integrates intelligent vision, intelligent voice, natural language understanding, real-time audio and video, and other 9 technologies and more than 20 AI capabilities, striving to create "deep intelligence" and maintain "high-tech configuration." The metaverse intelligent interaction engine of IConstruct Technology specifically includes four parts: the solution layer, MetaWorld SDK, editor, and content supply, reducing the threshold for creating metaverse scene play, allowing enterprises to try it out at low cost and quickly land specific play.
[0114] The virtual human dialogue communication module provides users with a vast number of virtual scenes that simulate and surpass real-world scenarios, mainly including office, learning, entertainment, and living scenarios. Users can use a joystick to select any scene and immerse themselves in it, engaging in dialogue with virtual characters in the scene, effectively practicing lip movements. The metaverse social scene provides a new way of communication and experience for hearing-impaired users, allowing them to immerse themselves in virtual space and communicate more freely, naturally, and effectively. The convenience of metaverse social activities anywhere, anytime also helps promote lip correction and lip movement practice among the deaf community, increasing their motivation for lip movement training, creating a virtuous cycle. The following will be a detailed introduction to the scene and user interaction to visualize the system.
[0115] A. Scene Introduction
[0116] Taking office and living scenarios as examples, the scene descriptions are as follows.
[0117] ① Office Scenario
[0118] The metaverse office scene is a virtual three-dimensional space where users can move freely and see, hear, and feel various elements in the virtual environment, such as buildings, desks, file cabinets, and more. The metaverse office scene supports online meetings and presentations, allowing users to hold meetings, present PPTs, videos, and more in the virtual space. Users can also organize and participate in team meetings in the metaverse office scene to discuss project progress, problem solving, strategy adjustment, and other related topics, and develop solutions together using virtual whiteboards and other tools.
[0119] ②Life scenarios
[0120] For example, in the metaverse shopping mall, users can see the mall's directional signs, stores, and people walking around. User A may be shopping in the mall and come to a store selling home appliances. User B, who works as a salesperson, greets them warmly. Users can communicate with each other to learn about the goods and negotiate prices.
[0121] B. User interaction explanation
[0122] Users can interact in the metaverse social scene. The following describes the interaction objects and interaction permissions.
[0123] a. Identity of interaction objects:
[0124] ① Virtual humans generated by the system interact with real humans
[0125] ② Virtual humans generated by the system interact with each other
[0126] ③ Real humans interact with each other
[0127] b. Operation permissions of interaction objects:
[0128] The system will determine the relationship between users and limit their operation permissions based on the relationship. First-level operation permissions are included in second-level operation permissions, and second-level operation permissions are included in third-level operation permissions.
[0129] ① If there is a blacklist relationship between users, enter the first-level operation permission, which specifically includes: users can see each other.
[0130] ② If users are strangers, enter the second-level operation permission, which specifically includes: before receiving the other party's reply, users can only say one sentence.
[0131] ③ If users follow each other, enter the third-level operation permission, which specifically includes: users can communicate anytime, anywhere, and interact with language and body.
[0132] Example 5
[0133] The functions of the user personal center module and the user's use of the user personal center module are described in detail in Example 5.
[0134] As Figure 5 shown, the user personal center module is equivalent to the user's personal center, in which the user can manage his own data and customize his image in the virtual community.
[0135] When the user selects the user personal center module, the system generates the user's personalized private space and builds a social scene that belongs to the user. Based on virtual human image generation and speech synthesis technology, the user can customize his virtual human image and voice to create a new social scene in the meta-universe.
[0136] In the user personal center module, the user can view and manage personal information. In addition to the basic functions of changing the nickname, changing the self-introduction, binding the contact information, and modifying the password, the module also includes customizing the virtual human image according to the user's image, viewing the user's practice duration, and viewing the user's practice effect. This module can help the user better understand his learning progress and learning situation, and manage the basic information of the account.
[0137] Based on the above needs, the invention uses a speaker face generation model based on wav21ip. The specific workflow is as follows:
[0138] For a speech segment, it is first converted into a mel spectrum form for easy processing, and encoded into an audio embedding through a multi-layer residual convolutional network. For a picture, down-sampling is done using a two-dimensional residual convolution, obtaining a picture embedding. For a video, each frame is processed in the same way as the picture. Transposed convolution (inverse convolution) is used as the decoder to reconstruct the picture.
[0139] A new loss function is added to the model. Specifically, an oral cavity discriminator is added to solve the problem of unsatisfactory oral cavity generation in the past model. Past models generally use L1 reconstruction loss as the loss function, and some models use discriminators to form GAN. Since the lips account for only about 4% of the entire face picture, the synchronization result of the past generation result on the lips is poor.
[0140] The oral cavity discriminator is composed of a pre-trained syncnet, which encodes the oral cavity and audio through two convolutional networks with the same structure, and evaluates the similarity between the encoded oral cavity and audio.
[0141] After investigation, the system of the present application finally selects Microsoft azure speech synthesis API. The speech synthesis technology of various companies is quite mature at present, and can achieve a very similar effect to a real person without being explicitly informed. At present, the more outstanding ones are Microsoft azure speech synthesis, Patsnap, Google Tacotron2, etc. In terms of the free version of the API, the speech synthesis of Patsnap has a mechanical sound and a pause, and the speech-based speaker face generation is prone to stuttering and jumping, thereby affecting the overall immersion experience of the project. By comparison, the Microsoft azure speech synthesis can produce better speech intonation, adjust the tone and pause according to the content, and bring a better interactive experience.
[0142] Those skilled in the art will readily understand that the above description is only preferred embodiments of the present application, and is not intended to limit the present application, and any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A lip-reading learning assistance training system based on the metaverse, characterized in that, include: Lip reading training module, virtual human dialogue and communication module, and user personal center module; The lip reading training module is used to store pre-collected standard lip-shape videos, establish a metaverse learning scenario, and enable users to train their lip reading through standard lip-shape videos in the metaverse learning scenario. It identifies the text read by the user's lips from the lip reading learning video when the user trains their lip reading through standard lip-shape videos, calculates the similarity between the text read by the user's lips and the text in the standard lip-shape video, and judges the user's lip reading training effect based on the similarity. The virtual human dialogue and communication module is used to establish a metaverse social scene, identify social text from the video of the user speaking in the metaverse social scene, convert the social text into audio during the dialogue process and combine it with the face to form a virtual human, so that the user can have a dialogue and communication with the virtual human in the metaverse social scene. The user personal center module is used to record and provide feedback on the user's lip reading training effect, and combines the user's audio with the face to form the user's virtual image, enabling the user to communicate with other users using the lip reading learning auxiliary training system in the metaverse social scene using the virtual image. The virtual human dialogue and communication module includes: a virtual human creation module and a dialogue robot. The virtual human formation module is used to call the lip reading training module's lip reading recognition model to identify social text from videos of users speaking in the metaverse social scene, input the social text into the dialogue robot, and then combine the dialogue robot's output response text into audio and combine it with the human face to form a virtual human. The virtual human creation module includes a speech synthesis module and an animation generation module. The speech synthesis module is used to synthesize audio from the text output by the chatbot using speech synthesis software. The animation generation module is used to combine audio with a human face using a speaking face generation model to form a virtual human; wherein, the speaking face generation model includes an encoder, a decoder, and a lip-sync discriminator, and the speaking face generation model is trained in the following manner: The sample speech segments are converted into Mel-spectrum form. The Mel-spectrum form of the sample speech segments is encoded into preprocessed audio through residual convolution in the encoder. The sample face image is downsampled through residual convolution in the encoder to obtain a preprocessed face image. The preprocessed audio and preprocessed face image are decoded through transposed convolution in the decoder to form a virtual person. The lip-sync discriminator encodes the lip-sync and audio of the virtual person through two convolutional networks. The goal is to minimize the error between the encoded lip-sync and the lip-sync in the preprocessed face image, and to minimize the error between the encoded audio and the preprocessed audio. The training continues until convergence, resulting in a trained speaking face generation model.
2. The lip-reading learning assistance training system based on the metaverse as described in claim 1, characterized in that, The lip-reading training module includes: a video preprocessing module, a lip-reading recognition module, and a feedback module. The video preprocessing module is used to store pre-collected standard lip-shape videos in multiple languages and to edit the standard lip-shape videos in each language into standard lip-shape videos in word mode and sentence mode. The lip reading recognition module is used to recognize the text read by the user's lips from lip reading learning videos when the user is training by lip reading using standard lip shape videos in word or sentence patterns in different languages; The feedback module is used to calculate the similarity between the text read by the user's lips and the text of the standard lip-reading video, and to judge the user's lip-reading training effect based on the similarity, and then feed the feedback to the user's personal center module.
3. The lip-reading learning assistance training system based on the metaverse as described in claim 2, characterized in that, The lip-reading training module includes a front-end feature extraction network and a back-end classification network, which are trained in the following way: Acquire face images and their actual lip language from video frames, extract the lip regions of the face images to form a ROI sequence, input the ROI sequence and the differenced ROI sequence into two branches of the front-end feature extraction network respectively, output the lip region features of the concatenated difference features, input the concatenated difference features of the lip region features into the back-end classification network, output the predicted characters, train until convergence with the goal of minimizing the error between the predicted characters and the actual lip language, and obtain the lip reading recognition model; The video frames are video frames in different languages, and finally, lip reading recognition models in different languages are obtained. The lip reading module is used to identify the text read by the user's lips from lip reading learning videos when the user is training by lip reading through standard lip-shape videos in word or sentence patterns in that language, using a lip reading recognition model for a specific language.
4. The lip-reading learning assistance training system based on the metaverse as described in claim 1, characterized in that, The chatbot is a personalized chatbot, and the personalization of the chatbot is achieved through the following methods: Collect dialogue texts from psychologists or teachers at schools for the hearing impaired. Before users interact with the chatbot, input the dialogue texts into ChatGPT, Wenxin Yiyan, WeChat Wormhole Assistant, PET, Bard, or MOSS chatbots to guide the chatbot to play the role of a psychologist or teacher at a school for the hearing impaired.
5. The lip-reading learning assistance training system based on the metaverse as described in claim 1, characterized in that, The lip-reading learning assistance system also includes: a metaverse scene creation module. The metaverse scene creation module is used to create a metaverse scene using Multispace or Baidu MetaStack. The lip reading training module is used to call the metaverse scene creation module to create a metaverse learning scene; The virtual human dialogue and communication module is used to call the metaverse scene creation module to create different metaverse social scenes; The virtual human formation module is used to identify social text in videos of users speaking in different metaverse social scenarios, input the social text into a dialogue robot, convert the dialogue robot's output response text into audio, and combine it with a human face to form virtual humans in different metaverse social scenarios, enabling users to communicate and interact with virtual humans in different metaverse social scenarios.
6. The lip-reading learning assistance training system based on the metaverse as described in claim 5, characterized in that, The user personal center module is used to store and manage the video data of the user using the lip-reading learning assistance training system to learn lip-reading, call the virtual human formation module to combine the user's audio with the face to form the user's virtual image, and call the metaverse social scene establishment module to establish the user's metaverse private space, so that the user can communicate with other users using the lip-reading learning assistance training system in the metaverse private space.
7. An application of a metaverse-based lip-reading learning assistance training system as described in any one of claims 1-6, characterized in that, The lip reading learning assistance training system is used to assist hearing-impaired people in learning lip reading. As users of the lip reading learning assistance training system, hearing-impaired people select standard lip shape videos from the lip reading training module to conduct lip reading training in the metaverse learning scenario, and judge the user's lip reading training effect by the similarity output by the lip reading training module. Users can select a virtual avatar from the virtual avatar dialogue module and engage in dialogue with the virtual avatar in the metaverse social scene; users can also customize their virtual avatar by selecting the user personal center module and engage in dialogue with other users using the lip-reading learning assistance training system in the metaverse social scene using the virtual avatar.
8. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the processing steps of the metaverse-based lip-reading learning aid training system according to any one of claims 1 to 6.
Citation Information
Patent Citations
Virtual-reality-based language training system and method
CN109064799A
Speaking face video generation method and device based on convolutional neural network
CN113378697A
Lip language recognition method and system
CN114973412A