Foreign language learning system and method using intelligent virtual human tutor, and computer program therefor
The system employs a virtual human tutor to analyze user data and provide personalized feedback, addressing the challenges of adult foreign language learning by offering affordable, time-flexible, and anxiety-reducing language practice.
Patent Information
- Application Number
- PCT/KR2024/018224
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-20
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-30
AI Technical Summary
Adults face significant challenges in learning foreign languages due to geographical and time constraints, high costs of professional tutors, and anxiety related to language interactions.
A system and method utilizing a virtual human tutor that analyzes user voice and image information to provide personalized foreign language learning content, allowing for real-time interaction and feedback without geographical or time limitations.
Enables effective and affordable foreign language learning by providing real-time customized feedback, reducing anxiety through simulated interactions, and accommodating learners at their convenience.
Smart Images

Figure KR2024018224_30052025_PF_FP_ABST
Abstract
Description
A foreign language learning system and method using an intelligent virtual human tutor and a computer program therefor
[0001] The embodiments relate to a foreign language learning system and method using a virtual human tutor, and a computer program therefor. More specifically, the embodiments relate to a technology for providing foreign language learning by analyzing a user's voice and video information using a virtual human tutor.
[0002] As we enter the global era, learning a foreign language is no longer an option but a necessity. The ability to speak a language other than one's native tongue plays a crucial role in enhancing individual and national competitiveness. However, learning a foreign language is not an easy task for anyone, and adults are known to find it even more challenging than children.
[0003] It's generally believed that younger children are more adept at language acquisition. It takes an average of 5,000 hours for a child to become fluent in their native language, and growing up in an environment completely immersed in it, they acquire it easily. In contrast, adults often struggle to acquire a new language without full exposure to it. Therefore, adult foreign language learning requires a different approach than that of children.
[0004] When learning a new language, it's essential to experience actual use and interaction with the language, such as living in that country, interacting with its people, and learning about its culture and customs along with the language itself. This also requires a broad understanding of the culture that speaks that language. While such interaction can be achieved through direct conversation with foreign friends, meeting them is often challenging. In particular, scheduling conflicts and geographical constraints make direct conversations sufficient for sufficient exposure to a foreign language challenging.
[0005] To overcome these challenges, many online services have been launched, enabling video language exchange and learning with foreign tutors. These services coordinate the schedules of foreign tutors and language learners online, enabling language interaction and supporting language learning. Online video chat services are considered effective tools for foreign language learning.
[0006] However, although learning a foreign language online is possible, there are still time and geographical constraints (time differences, scheduling difficulties, etc.) with foreign tutors, and the high cost of lessons and services from professional foreign tutors.
[0007] Additionally, learners with limited foreign language skills may feel anxious when talking with foreign tutors, and this anxiety, which commonly occurs during foreign language learning, has the problem of negatively affecting language learning.
[0008] In this way, learning a foreign language is one of the important skills for modern society and humanity, and the skills to learn a foreign language on one's own are in demand.
[0009] According to one aspect of the present invention, a foreign language learning system and method using a virtual human tutor, which analyzes a user's voice information and image information using the virtual human tutor to provide content for foreign language learning, and a computer program therefor can be provided.
[0010] A foreign language learning method according to one aspect of the present invention comprises the steps of: providing content information including a profile of a virtual human tutor, a category of a conversation topic to be learned, and a virtual environment corresponding to the category; receiving voice information and video information of a user; regenerating the virtual environment by changing a field of view or an image of the virtual environment based on the video information; generating response data of the virtual human tutor corresponding to the voice information of the user; generating voice data and motion data of the virtual human tutor based on the response data of the virtual human tutor; and rendering the voice data and motion data of the virtual human tutor to provide content.
[0011] In one embodiment, the step of changing the field of view or image for the virtual environment includes the step of extracting the user's gesture information from the image information; and the step of generating the learning comprehension information by comparing the user's gesture information with previously labeled learning gesture information.
[0012] In one embodiment, the step of generating voice data and motion data of the virtual human tutor includes the step of generating voice data of the virtual human tutor by adjusting the speech speed of the virtual human tutor based on the learning comprehension information.
[0013] In one embodiment, the step of generating response data of the virtual human tutor includes the step of extracting the user's foreign language level based on the voice information; and the step of generating response data of the virtual human tutor by taking the foreign language level into consideration.
[0014] A computer program according to one aspect of the present invention may be stored in a computer-readable recording medium so as to be combined with hardware and execute a foreign language learning method according to any one of claims 1 to 4.
[0015] A foreign language learning system according to one aspect of the present invention comprises: a content information providing unit that provides content information including a profile of a virtual human tutor, a category of a conversation topic to be learned, and a virtual environment corresponding to the category; a receiving unit that receives voice information and image information of a user; a content generating unit that regenerates the virtual environment by changing a field of view or an image of the virtual environment based on the image information, generates response data of the virtual human tutor corresponding to the voice information of the user, and generates voice data and motion data of the virtual human tutor based on the response data of the virtual human tutor; and a content providing module that provides content by rendering the voice data and motion data of the virtual human tutor.
[0016] According to a digital content production system and method according to one aspect of the present invention, a user can learn a foreign language regardless of the time and place he or she desires, can learn at a lower price compared to the service cost of an actual foreign tutor, and can provide real-time customized feedback to the user through artificial intelligence analysis.
[0017] Additionally, users with limited foreign language skills can improve their foreign language skills by speaking and practicing foreign languages with a virtual human tutor.
[0018] Figure 1 is a schematic block diagram of a foreign language learning system according to one embodiment.
[0019] Figure 2 is a schematic block diagram showing the hardware configuration of a foreign language learning system according to one embodiment.
[0020] Figure 3 is a flowchart showing each step of a foreign language learning method according to one embodiment.
[0021] FIGS. 4A to 4D are exemplary images showing a user interface (UI) on a user device according to a foreign language learning method according to one embodiment.
[0022] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0023] Figure 1 is a schematic block diagram of a foreign language learning system according to one embodiment.
[0024] Referring to FIG. 1, the foreign language learning system (2) according to the present embodiment is configured to communicate with and operate a user device (1), thereby providing content information including a virtual human tutor, a category of a conversation topic, and a virtual environment corresponding to the category to the user device (1), and to render voice data and motion data of the virtual human tutor to provide content.
[0025] For the above operation, the foreign language learning system (2) is configured to be able to communicate with the user device (1), microphone (3), and photographing device (4) via a wired and / or wireless network. At this time, the wired and / or wireless networks include LAN (Local Area Network), MAN (Metropolitan Area Network), GSM (Global System for Mobile Network), EDGE (Enhanced Data GSM Environment), HSDPA (High Speed Downlink Packet Access), W-CDMA (Wideband Code Division Multiple Access), CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), Bluetooth, Zigbee, Wi-Fi, VoIP (Voice over Internet Protocol), LTE Advanced, IEEE802.16m, WirelessMAN-Advanced, HSPA+, 3GPP Long Term Evolution (LTE), Mobile WiMAX (IEEE 802.16e), UMB (formerly EV-DO Rev. C), Flash-OFDM, iBurst and MBWA (IEEE 802.20) systems, HIPERMAN, Beam-Division Multiple Access (BDMA), Wi-MAX (World Interoperability for Microwave Access) and may refer to a communication network using one or more communication methods selected from the group consisting of ultrasonic communication, but is not limited thereto.
[0026] A user device (1) is a device used by a user, such as a student or the general public, who wishes to learn a foreign language through the functions provided by the foreign language learning system. A user can access the foreign language learning system (2) and use the services provided by the foreign language learning system (2) by executing a predetermined application (or app) on his / her user device (1).
[0027] The number and shape of the user devices (1) illustrated in FIG. 1 are merely exemplary. For example, although the user device (1) in FIG. 1 is illustrated in the form of an XR (Extended Reality) device, in other embodiments, the user device (1) may be implemented in the form of any computing device, such as a mobile communication terminal such as a smartphone, a notebook computer, a personal computer, a personal digital assistant (PDA), a tablet, a set-top box for IPTV (Internet Protocol Television), etc.
[0028] Meanwhile, in the drawings attached to this specification, the foreign language learning system (2) is depicted as a separate device from the user device (1), but this is exemplary, and depending on the embodiment, the foreign language learning system (2) may be implemented in the form of a software application stored and executed on the user device (1).
[0029] The foreign language learning system (2) includes a database (DB) (20), a content management module (21), a voice recognition module (22), a motion recognition module (23), and a content provision module (24). In one embodiment, the foreign language learning system (2) further includes an output module (25). Each module (21-25) may include one or more functional units. Meanwhile, each of these units or modules (21-25) may be realized at least partially using the hardware (200) of the foreign language learning system (2).
[0030] That is, the foreign language learning system (2) according to the embodiments and each part or module (21-25) included therein may have aspects that are entirely hardware, or partially hardware and partially software. For example, each part and module (21-25) of the foreign language learning system (2) may collectively refer to hardware and related software for processing data of a specific format and content or / and exchanging data via electronic communication. In this specification, terms such as "part," "module," "device," "terminal," "server," or "system" are intended to refer to a combination of hardware and software driven by the hardware. For example, the hardware may be a data processing device including a CPU or other processor. In addition, the software driven by the hardware may refer to a running process, an object, an executable, a thread of execution, a program, etc.
[0031] In addition, each element constituting the foreign language learning system (2) is not necessarily intended to refer to a separate device that is physically distinct from one another. That is, each part and module (21-25) of the foreign language learning system (2) illustrated in Fig. 1 is merely a functional division of the hardware constituting the foreign language learning system (2) according to the operations performed by the hardware, and each part does not necessarily have to be provided independently from one another. Of course, depending on the embodiment, one or more of the above-described parts and modules (21-25) may be implemented as a separate device that is physically distinct from one another.
[0032] The user device (1) may be implemented as an XR device that allows the user to view and converse with a virtual human tutor. For example, commercial products such as Apple's Apple Vision Pro may be used as the XR device, but the present invention is not limited thereto.
[0033] The virtual human tutor generated by the foreign language learning system (2) may be provided in a form that reflects the user's physical space using spatial computing. Spatial computing can be utilized more effectively when the user device (1) is an XR device.
[0034] The user device (1) can receive content information including a profile of a virtual human tutor, a category of a conversation topic to be learned, and a virtual environment corresponding to the category.
[0035] The user device (1) can receive content in which voice data and motion data of a virtual human tutor are rendered based on response data of the virtual human tutor.
[0036] The microphone (3) acquires the voice of the speaking user.
[0037] Here, the microphone (3) can acquire the voice of the user conversing with the virtual human tutor.
[0038] Additionally, the microphone (3) can acquire voice information of the user's selection of a virtual human tutor or a category of conversation topic and virtual environment.
[0039] The microphone (3) transmits the acquired user's voice to the foreign language learning system (2).
[0040] The microphone (3) is implemented as a device capable of recording sound, and may be configured separately from the user device (1), or may be a microphone built into the user device (1).
[0041] The photographing device (4) photographs the user's movements.
[0042] The photographing device (4) can photograph the user's face and body and transmit the photographed image of the user as gesture information to the foreign language learning system (2).
[0043] The photographing device (4) is implemented as a device capable of photographing, and may be configured separately from the user device (1), or may be a camera built into the user device (1).
[0044] The foreign language learning system (2) provides content to users by rendering voice data and movement data of a pre-generated virtual human tutor.
[0045] DB (20) can store content information including a profile of a virtual human tutor, a category of a conversation topic to be learned, and a virtual environment corresponding to the category.
[0046] Here, the profile of the virtual human tutor may include at least one of country, language, name, appearance gender, voice, personality, and interests.
[0047] Additionally, the categories of conversation topics can include at least one of daily life, hobbies, history, travel, culture, exercise, self-improvement, food, career, news, event information, economy, and politics.
[0048] Additionally, the virtual environment can be implemented with a metaverse and 3D computer graphics, and can be composed of specific places reflecting the culture of each country, and can include at least one of each country's famous landmarks, nature, historical sites, museums and art galleries, restaurants, schools, and government offices, which are specific places reflecting the culture of each country.
[0049] DB (20) can store response data, voice data, and motion data of the generated virtual human tutor.
[0050] DB (20) can store the user's voice information and video information.
[0051] DB (20) can store a video of a user or avatar conversing with a virtual human tutor using a 3D virtual camera.
[0052] That is, DB (20) stores the conversations and learning between the user and the virtual human tutor, so that the user can view them at any time.
[0053] The content management module (21) may include a content information provision unit (211) and a content creation unit (212).
[0054] The content information provision unit (211) provides content information including a virtual human tutor, a category of a conversation topic to be learned, and a virtual environment corresponding to the category to the user through the user terminal (1).
[0055] Here, the virtual human tutor can be visualized as a virtual human or virtual character with human-like intelligence and behavior created through three-dimensional computer graphics based on artificial intelligence (AI) technology.
[0056] The virtual human tutor is equipped with a large language model (LLM) such as GPT, can think for itself, support multiple languages, and can have natural conversations with users through text-to-speech (TTS).
[0057] Additionally, the virtual human tutor can be trained to include information about language and culture, and can perform conversations and simulate situations relevant to the user and their cultural background.
[0058] The virtual human tutor can detect the user's emotional state and situation based on the user's video information, and recognize multimodal data such as the user's visual, auditory, and biometric information.
[0059] The content generation unit (212) can regenerate a virtual environment by changing the field of view or image of the virtual environment based on the user's image information.
[0060] The content generation unit (212) can generate response data of a virtual human tutor corresponding to the user's voice information, extract the user's foreign language level based on the voice information, and generate response data of the virtual human tutor by taking the foreign language level into consideration.
[0061] The content generation unit (212) can generate learning comprehension information by comparing pre-labeled learning gesture information with the user's gesture information, and can generate voice data of the virtual human tutor by adjusting the speech speed of the virtual human tutor based on the learning comprehension information.
[0062] In one embodiment, the content generation unit (212) can generate an avatar for the user.
[0063] Here, the content generation unit (212) can reflect the user's gesture information included in the user's video information to the user's avatar in real time.
[0064] The voice recognition module (22) may include a receiving unit (221) and a voice recognition unit (222).
[0065] The receiver (221) can receive the user's voice information transmitted through the microphone (3).
[0066] Here, the user's voice information may be information for conversation with a virtual human tutor, and may be voice information for the user to select a category of a virtual human tutor or conversation topic and a virtual environment.
[0067] The voice recognition unit (222) can recognize words used by the user based on the received voice information of the user.
[0068] The voice recognition unit (222) can analyze the user's voice through voice recognition and natural language processing technology to recognize words included in the voice information.
[0069] Additionally, the voice recognition unit (222) can further recognize and extract information about the user's foreign language pronunciation and intonation.
[0070] The motion recognition module (22) may include a receiving unit (231) and a gesture recognition unit (232).
[0071] The receiving unit (231) can receive an image of the user captured through the capturing device (4).
[0072] The gesture recognition unit (232) can extract the user's gesture information from the received user's image.
[0073] Here, gesture information may include right or left "sleeve tilt* information based on the front of the face or movement information in the up-down direction of the face based on the front of the face.
[0074] Additionally, gesture information may further include facial movement information including facial expression information, eyebrow movement information, pupil position information, and lip movement information, and body movement information including hand movement information and foot movement information.
[0075] The gesture recognition unit (232) can recognize gesture information through real-time human pose estimation technology such as Mediapipe, and there may be room for the technology to change as technology advances.
[0076] The content provision module (24) renders voice data and motion data of the generated virtual human tutor and provides the content to the user through the user device (1).
[0077] The content provision module (24) can provide the regenerated virtual environment to the user device (1) in real time.
[0078] The content provision module (24) can provide content to the user device (1) by reflecting voice data and motion data to the virtual human tutor in real time.
[0079] The content provision module (24) can provide improvements in foreign language skills by comparing the user with other users based on the content of the conversation and the recorded video with the virtual human tutor.
[0080] The output module (25) can output content generated by rendering voice data and motion data.
[0081] Here, the output module (25) can transmit to the user device (1) or a separate external server (not shown).
[0082] The output module (25) can transmit the conversation and learning between the user and the virtual human tutor to the user device (1) or a separate external server.
[0083] The output module (25) can output a personalized learning plan based on the user's learning comprehension information and foreign language level.
[0084] The output module (25) can output comprehensive analysis data and reports that analyze the user's voice information and gesture information to support the user's improvement of foreign language skills.
[0085] In one embodiment, the output module (25) can output an analysis result comparing the extracted user's foreign language pronunciation and intonation with the virtual human tutor's foreign language pronunciation and intonation.
[0086] The output module (25) can output a certificate that can authenticate the user's foreign language ability and history evaluated through a virtual human tutor.
[0087] Figure 2 is a schematic block diagram showing the hardware configuration of a foreign language learning system according to one embodiment.
[0088] Referring to FIG. 2, a foreign language learning system (2) according to embodiments is implemented in the form of a computing device including hardware (200), and may include a memory (210), a processor (220), a communication module (230), and an input / output unit (240).
[0089] The memory (210) is a non-transitory computer-readable recording medium and may include a permanent mass storage device such as a random access memory (RAM), a read only memory (ROM), a disk drive, a solid state drive (SSD), a flash memory, etc. Here, the non-permanent mass storage device such as a ROM, an SSD, a flash memory, a disk drive, etc. may be included in the above-described device or server as a separate permanent storage device distinct from the memory (210).
[0090] Additionally, the memory (210) may store an operating system and at least one program code (e.g., code for a security module or an application installed to provide a specific service). These software components may be loaded from a computer-readable recording medium separate from the memory (210). This separate computer-readable recording medium may include a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, or a memory card.
[0091] In another embodiment, the software components may be loaded into the memory (210) via a communication module (230) rather than a computer-readable recording medium. For example, at least one program may be loaded into the memory (210) based on a computer program that is installed by files provided over a network by developers or a file distribution system (e.g., an application store service server) that distributes installation files for applications.
[0092] The processor (220) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (220) by the memory (210) or the communication module (230). For example, the processor (220) may be configured to execute instructions received according to program code stored in a storage device such as the memory (210).
[0093] The communication module (230) may provide a function for the foreign language learning system (2) to communicate with the user device (1), microphone (3), photographing device (4), etc. via a network. In addition, the communication module (230) may provide a function for the foreign language learning system (2) to communicate with one or more other devices via a wired and / or wireless network. That is, the communication module (230) is a part that realizes each functional module described above with reference to FIG. 1 by having its function controlled by a processor (220) referencing a memory (210).
[0094] The input / output unit (240) may be a means for interfacing with an external input / output device (not shown). For example, external input devices may include devices such as a keyboard, mouse, microphone, camera, etc., and external output devices may include devices such as a display, speaker, haptic feedback device, etc. As another example, the input / output unit (240) may be a means for interfacing with a device that integrates input and output functions, such as a touchscreen.
[0095] In addition, in other embodiments, the foreign language learning system (2) may include more hardware components than those illustrated in FIG. 2, depending on the nature of the device to which it is applied. For example, when the foreign language learning system (2) is applied to a terminal device used by a user, it may be implemented to include at least some of the above-described input / output devices, or may further include other components such as a transceiver, a GPS (Global Positioning System) module, a camera, various sensors, a DB, etc. As a more specific example, when the terminal device is a smartphone, it may be implemented to further include various components that are generally included in a smartphone, such as an acceleration sensor or a gyro sensor, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.
[0096] However, the components and forms of the computing device described in this specification are merely exemplary, and the configuration of the computing device in which the foreign language learning system (2) is implemented may differ from that described in this specification depending on the adoption of other known technologies or the development of future information and communication technologies.
[0097] The foreign language learning method described below can be performed by a foreign language learning system (2) implemented in the form of a computing device including the hardware (200) configuration described above with reference to FIG. 2. For example, the foreign language learning method can be provided to a user in the form of a service based on at least one of an application, software, or other program operating on a user device and / or a server.
[0098] FIG. 3 is a flowchart showing each step of a foreign language learning method according to one embodiment, and FIGS. 4a to 4d are exemplary images showing a user interface (UI) on a user device by a foreign language learning method according to one embodiment.
[0099] When a user is connected to a foreign language learning system (2), the content information provision unit (211) provides a profile of a virtual human tutor to the user device (1) so that the user can preview and select a virtual human tutor (S11).
[0100] Referring to FIG. 4a, the content information provision unit (211) can provide a profile of a virtual human tutor including the country, language, and name of the virtual human tutor.
[0101] Here, the profile of the virtual human tutor is depicted as including country, language, and name, but the profile of the virtual human tutor may include at least one of country, language, name, appearance gender, voice, personality, and interests.
[0102] In addition, the content information provision unit (211) can provide a function that can be customized according to the user's tendencies and learning requirements, thereby creating and providing a customized virtual human tutor.
[0103] Here, the content information provision unit (211) can receive optional information from the user regarding the desired country, language, gender, personality, interests, etc., and create a virtual human tutor customized for the user.
[0104] The user selects a virtual human tutor with whom he or she wishes to converse and learn by checking the profile of the virtual human tutor through the user device (1) (S12).
[0105] A user can select a profile of a virtual human tutor in a virtual environment through a user device (1) such as an XR device.
[0106] The user may select a virtual human tutor by voice or gesture, or may select a virtual human tutor by using a joystick connected to the user device (1).
[0107] The content information provision unit (211) provides categories and virtual environments for conversation topics to be learned (S13).
[0108] Referring to FIG. 4b, the content information provision unit (211) can provide categories of conversation topics including daily life, hobbies, history, and travel to the user device (1).
[0109] Here, the categories of conversation topics are shown as including daily life, hobbies, history, and travel, but the categories of conversation topics can include at least one of daily life, hobbies, history, travel, culture, exercise, self-development, food, job, news, event information, economy, and politics.
[0110] Additionally, the content information provision unit (211) can provide a virtual environment corresponding to the category of the conversation topic.
[0111] Here, the virtual environment may be composed of specific places reflecting the culture of each country, and may include at least one of each country's famous landmarks, nature, historical sites, museums and art galleries, restaurants, schools, and government offices.
[0112] The content information provision unit (211) can automatically provide the user with a virtual environment corresponding to the category of the conversation topic.
[0113] For example, if a user selects travel as a category of conversation topic, the content information provision unit (211) can automatically match a virtual environment to the travel location and provide it to the user.
[0114] The user selects a category of a topic and a virtual environment in which he or she wishes to converse with a virtual human tutor through the user device (1) (S14).
[0115] Users can select categories of conversation topics and virtual environments through user devices (1) such as XR devices.
[0116] The user may select a category of conversation topic and a virtual environment by voice or gesture, or may select a category of conversation topic and a virtual environment by using a joystick connected to the user device (1).
[0117] That is, referring to FIG. 4c, the content provision module (24) can generate content based on a virtual human tutor (610) and virtual environment (600) selected by the user and provide the content through the user device (1).
[0118] For example, the content information provider (211) may provide a virtual human tutor (610) and the Eiffel Tower and the Louvre Museum as a virtual environment (600).
[0119] When a user and a virtual human tutor start a conversation based on the generated content, the filming device (4) transmits the user's video information, which was captured while the user is speaking, to the foreign language learning system (2) (S15).
[0120] The receiving unit (231) of the motion recognition module (23) can receive the user's image information.
[0121] The content creation unit (212) creates learning comprehension information using the user's video information (S16).
[0122] Specifically, the gesture recognition unit (232) can extract the user's gesture information from the user's image information.
[0123] Here, gesture information may include right or left "sleeve tilt* information based on the front of the face or movement information in the up-down direction of the face based on the front of the face.
[0124] Additionally, gesture information may further include facial movement information including facial expression information, eyebrow movement information, pupil position information, and lip movement information, and body movement information including hand movement information and foot movement information.
[0125] In one embodiment, the content generation unit (212) can generate learning comprehension information by comparing pre-labeled learning gesture information with user gesture information.
[0126] Here, the labeled learning gesture information represents data classified by collecting gesture data of multiple learners and labeling learning comprehension information of each of the multiple gesture data.
[0127] Learning comprehension information represents information about whether the virtual human tutor understands the words or sentences spoken.
[0128] In another embodiment, the content generation unit (212) may generate learning comprehension information with a relatively low level by recognizing that the user does not understand when the tilt information included in the user's gesture information exceeds a threshold or the movement information falls outside a reference range.
[0129] In another embodiment, the content generation unit (212) can input the user's gesture information into a pre-learned artificial intelligence-based analysis model to generate learning comprehension information.
[0130] Here, the analysis model can be trained using pre-labeled learning gesture information.
[0131] That is, the content generation unit (212) can generate learning comprehension information corresponding to the user's gesture information.
[0132] The content generation unit (212) regenerates the virtual environment by changing the field of view or image of the virtual environment based on the user's image information (S17).
[0133] Specifically, the content generation unit (212) can regenerate the virtual environment in real time by changing the field of view or image of the virtual environment in response to learning comprehension information generated from the user's image information.
[0134] For example, when the user's learning comprehension information is relatively high, the content creation unit (212) can create a virtual environment in which the image can be viewed as a whole, as in Fig. 4c.
[0135] The content generation unit (212) can regenerate an expanded virtual environment by reducing the field of view of the virtual environment when the user's learning comprehension information is relatively low.
[0136] In this way, the content generation unit (212) can regenerate the virtual environment by adjusting the field of view of the virtual environment in response to the user's learning comprehension information and provide the regenerated virtual environment to the user device (1).
[0137] In addition, the content generation unit (212) can regenerate the virtual environment by enlarging or reducing the mouth shape of the virtual human tutor so that the mouth shape of the virtual human tutor can be seen in response to the user's learning comprehension information.
[0138] In addition, referring to FIG. 4d, the content generation unit (212) can regenerate the virtual environment by changing the image of the virtual environment to a simple image when the user's learning comprehension information is relatively low.
[0139] When the user speaks, the microphone (3) provides the user's voice information to the foreign language learning system (2) (S18).
[0140] The receiving unit (221) of the voice recognition module (22) can receive the user's voice information.
[0141] The content generation unit (212) generates response data of a virtual human tutor corresponding to the user's voice information (S19).
[0142] Specifically, the voice recognition unit (222) can recognize words used by the user based on the user's voice information.
[0143] In one embodiment, the content generation unit (212) can extract the user's foreign language level based on recognized words using pre-labeled words.
[0144] Here, the labeled words indicate the foreign language level labeled for each word.
[0145] The content generation unit (212) can extract the user's foreign language level by matching pre-labeled words with recognized words.
[0146] Here, if there are multiple recognized words, the average foreign language level of each word can be extracted as the user's foreign language level.
[0147] In another embodiment, the content generation unit (212) can input the user's voice information into a pre-trained artificial intelligence-based analysis model to output a foreign language level.
[0148] Here, the analysis model can be trained by collecting pre-labeled words.
[0149] The content generation unit (212) can generate response data of a virtual human tutor by taking into account the user's foreign language level.
[0150] In this way, the content generation unit (212) can generate response data of a virtual human tutor composed of words that the user can understand, taking into account the user's foreign language level.
[0151] In one embodiment, the foreign language learning system (2) depicts the user's video information reception (S15) and the user's voice information reception (S18) as separate steps, but they can be performed simultaneously and in real time.
[0152] The content generation unit (212) generates voice data and motion data of a virtual human tutor based on the response data (S20).
[0153] The content generation unit (212) can generate voice data of a virtual human tutor by adjusting the speech speed of the virtual human tutor based on the previously generated learning comprehension information.
[0154] The content generation unit (212) can generate voice data of the virtual human tutor so that the user can understand it by adjusting the speech speed of the virtual human tutor to be slow or fast in response to the learning comprehension information.
[0155] In addition, the content generation unit (212) can automatically generate voice data and motion data of a virtual human tutor by inputting response data into a pre-learned generative artificial intelligence model.
[0156] The content generation unit (212) can generate voice data according to the voice of a virtual human tutor selected by the user, and can generate motion data corresponding to the response data.
[0157] The content provision module (24) renders the generated voice data and motion data and provides the content to the user device (1) (S21).
[0158] Here, the content provision module (24) can render the generated voice data and motion data to a virtual human tutor to provide content to the user device (1) in real time.
[0159] Referring again to FIG. 4d, the content provision module (24) can reflect the generated voice data and motion data to the virtual human tutor.
[0160] Additionally, the content provision module (24) can provide improvements in foreign language proficiency by comparing the user with other users based on the content of the conversation and the recorded video with the virtual human tutor.
[0161] That is, in addition, the content provision module (24) can provide content including a virtual human tutor that reflects the regenerated virtual environment in real time, moves according to generated motion data, and speaks according to voice data to the user device (1) in real time.
[0162]
[0163] The operations of the foreign language learning method according to the embodiments described above can be implemented at least partially as a computer program and recorded on a computer-readable recording medium. The computer-readable recording medium on which the program for implementing the operations of the foreign language learning method according to the embodiments is recorded includes all types of recording devices that store data that can be read by a computer. Examples of the computer-readable recording medium include ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. In addition, the computer-readable recording medium can be distributed across network-connected computer systems, so that the computer-readable code can be stored and executed in a distributed manner. In addition, the functional programs, codes, and code segments for implementing the present embodiment can be easily understood by those skilled in the art to which the present embodiment belongs.
[0164] While the present invention has been described above with reference to the embodiments illustrated in the drawings, these are merely exemplary, and those skilled in the art will appreciate that various modifications and variations of the embodiments are possible. However, such modifications should be considered within the technical protection scope of the present invention. Therefore, the true technical protection scope of the present invention should be determined by the technical spirit of the appended claims.
[0165] The embodiments relate to a foreign language learning system and method using a virtual human tutor, and a computer program therefor. More specifically, the embodiments relate to a technology for providing foreign language learning by analyzing a user's voice and video information using a virtual human tutor.
Claims
1. A step of providing content information including a profile of a virtual human tutor, a category of a conversation topic to be learned, and a virtual environment corresponding to the category; A step of receiving user's voice information and image information; A step of regenerating the virtual environment by changing the field of view or image of the virtual environment based on the image information; A step of generating response data of the virtual human tutor corresponding to the voice information of the user; A step of generating voice data and motion data of the virtual human tutor based on the response data of the virtual human tutor; and A step of providing content by rendering voice data and motion data of the virtual human tutor; A method for learning foreign languages using a virtual human tutor.
2. In paragraph 1, The step of changing the field of view or image for the above virtual environment is: A step of extracting the user's gesture information from the above image information; A method for learning a foreign language using a virtual human tutor, comprising: a step of generating learning comprehension information by comparing labeled learning gesture information with gesture information of the user; 3. In paragraph 2, The step of generating voice data and motion data of the above virtual human tutor is: A method for learning a foreign language using a virtual human tutor, comprising the step of generating voice data of the virtual human tutor by adjusting the speech speed of the virtual human tutor based on the learning comprehension information.
4. In paragraph 1, The step of generating response data of the above virtual human tutor is: A step of extracting the user's foreign language level based on the above voice information; and A method for learning a foreign language using a virtual human tutor, comprising: a step of generating response data of the virtual human tutor by considering the foreign language level.
5. A computer program stored in a computer-readable recording medium that is combined with hardware and executes a foreign language learning method according to any one of claims 1 to 4.
6. A content information provision unit that provides content information including a profile of a virtual human tutor, a category of a conversation topic to be learned, and a virtual environment corresponding to the above category; A receiving unit that receives user voice information and image information; A content generation unit that regenerates the virtual environment by changing the field of view or image of the virtual environment based on the image information, generates response data of the virtual human tutor corresponding to the voice information of the user, and generates voice data and motion data of the virtual human tutor based on the response data of the virtual human tutor; and A content providing module that provides content by rendering voice data and motion data of the virtual human tutor; A foreign language learning system using a virtual human tutor.
Citation Information
Patent Citations
Low dielectric high heat dissipation film composition for 5G FCCL and manufacturing method thereof
KR1020240065705A
System for providing seasonal worker total management service
KR102596482B1
Upconversion multi-color-emitting polymer composite, transparent display including the same and method for manufacturing the same
KR102762614B1
Educational teaching system and method utilizing interactive avatars with learning manager and authoring manager functions
US20170206797A1
KR20190078294A