Information processing apparatus, information processing method, and information processing program

The information processing device uses a trained generative model to create an avatar for language learning, addressing user anxiety and enhancing motivation by providing a supportive conversational environment and feedback.

JP2025121688AInactive Publication Date: 2025-08-20SOFTBANK GROUP CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024017300
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-07
Publication Date
2025-08-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Conventional dialogue systems face challenges in motivating users to learn a language due to nervousness and anxiety about conversing with others, using incorrect grammar or vocabulary, and repetitive questioning.

Method used

An information processing device generates an avatar using a trained generative model to engage in conversational sentences tailored to the user's learning conditions, providing a comfortable language learning environment by receiving inputs, generating responses, and outputting them in the desired language.

Benefits of technology

The system enhances user motivation by alleviating anxiety, allowing users to practice without fear of judgment and offering feedback for review, thus improving language learning outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025121688000001_ABST
    Figure 2025121688000001_ABST
Patent Text Reader

Abstract

To provide a language learning with a lower psychological hurdle.SOLUTION: An information processing apparatus comprises: a receiving unit 13a that receives a user's selection; a generation unit 13c that, on the basis of the selection received by the receiving unit 13a, generates an avatar controlled by a control unit comprising a generation model subjected in advance to language learning to generate an answer to an input prompt for a conversion sentence in a specific language; and an output control unit 13e that controls the avatar generated by the generation unit 13c to perform the conversion sentence in the specific language for the user.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] BACKGROUND ART Conventionally, a dialogue system is known that generates an answer to a question input by a user using a language model (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Special Publication No. 2022-503838 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional dialogue systems have difficulty motivating users to learn a language. For example, when learning a foreign language, users may feel nervous about conversing with others. They may also feel anxious about using the same grammar or vocabulary incorrectly or about being asked the same questions over and over again.

[0005] As described above, conventional technologies make it difficult for users to study a language with peace of mind, and language learning also poses high psychological hurdles, making it difficult to increase users' motivation to study a language.

[0006] The present application aims to increase users' motivation to learn languages. [Means for solving the problem]

[0007] The information processing device of the present application includes a reception unit that receives a user's selection, a generation unit that generates an avatar controlled by a control device equipped with a generation model that has undergone language training in advance to generate conversational sentences in a specific language to generate responses to input prompts based on the selection received by the reception unit, and an output control unit that controls the avatar generated by the generation unit to speak the conversational sentences in the specific language to the user. [Effects of the Invention]

[0008] According to one aspect of the embodiment, it is possible to improve a user's motivation to learn a language. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the overall configuration of a system according to the first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of language learning in the information processing device according to the first embodiment. [Figure 3] FIG. 3 is a functional block diagram showing the functional configuration of the information processing device 10 according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating the user information DB. [Figure 5] FIG. 5 is a diagram illustrating the situation DB. [Figure 6] FIG. 6 shows an example of a screen display on which the user selects learning conditions in the first embodiment. [Figure 7] FIG. 7 is an example of a display of an avatar generated in the first embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of a conversation between a user and an avatar according to the first embodiment. [Figure 9] FIG. 9 shows an example of a display of a feedback screen for a conversation between a user and a generative model according to the first embodiment. [Figure 10] FIG. 10 is a flowchart showing an example of the flow of processing in the information processing device according to the first embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of the overall configuration of a system according to the second embodiment. [Figure 12] FIG. 12 is a functional block diagram showing the functional configuration of the information processing device 10 according to the second embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of order information stored in the information processing device according to the second embodiment. [Figure 14] FIG. 14 is an example of a display of an order screen for user selection according to the second embodiment. [Figure 15] FIG. 15 is a flowchart showing an example of the flow of processing in the information processing device according to the second embodiment. [Figure 16] FIG. 16 is a diagram illustrating an example of the overall configuration of a system according to the third embodiment. [Figure 17] FIG. 17 is an example of a diagram illustrating avatar relearning according to the third embodiment. [Figure 18] FIG. 18 is a diagram illustrating a user information DB according to the fourth embodiment. [Figure 19] FIG. 19 is a diagram illustrating a situation DB according to the fourth embodiment. [Figure 20] FIG. 20 shows an example of a display screen for user selection according to the fourth embodiment. [Figure 21] FIG. 21 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of the information processing device. DETAILED DESCRIPTION OF THE INVENTION

[0010] The present invention will be described below through embodiments, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.

[0011] <1. Embodiment> (1-1. Overview of the First Embodiment) In the first embodiment, a conversation between a user and a generated avatar will be described. FIG. 1 is a diagram illustrating an example of the overall configuration of a system according to the first embodiment. As shown in FIG. 1, the system according to the first embodiment includes an information processing device 10, a terminal device 20, and a language information database 30.

[0012] The information processing device 10 is an example of a computer that generates an avatar and a response sentence based on the learning conditions of a user who is learning a language, outputs various information to a terminal device, and carries out a conversation with the user.

[0013] The terminal device 20 is an example of a terminal device used by a user. For example, the terminal device 20 is a mobile phone, a smartphone, a tablet terminal, a personal computer, etc. The terminal device 20 also has an audio input device (e.g., a microphone) that receives audio input, and an audio output device (e.g., a speaker) that outputs audio.

[0014] The language information database 30 is a database that stores various data related to language. For example, the language information database 30 includes various languages such as English, Japanese, and Chinese, and is used to extract commonalities in grammar rules and vocabulary using natural language processing technology. The language information database 30 is also used to process and analyze language data using AI (Artificial Intelligence). Note that the number of terminal devices 20 and language information databases 30 installed in the system is not limited to the number shown in FIG. 1.

[0015] In this embodiment, the information processing device 10 learns a language model based on a data source acquired from the language information database 30. The language model here can be any of various well-known models such as generative AI. Algorithms such as GAN (Generative Adversarial Networks) and VAE (Variational Autoencoder) that use neural networks can be used for this language model.

[0016] (General language learning problems) Typical language learning applications and systems have difficulty motivating users to learn a language. For example, if a user wants to learn English, they can either meet face-to-face with an English teacher or converse with the teacher using an application or the Internet. However, users must study face-to-face without confidence in their English conversation skills, which can make them feel nervous about conversing with others. Users may also feel anxious about using incorrect English words or grammar, or worry that incorrect usage could prevent them from conveying their intended message. Furthermore, users may feel anxious about repeatedly asking the same questions to their learning partners. As described above, users' feelings of tension and anxiety contribute to a decline in their motivation to learn.

[0017] Therefore, the information processing device 10 according to the first embodiment provides an information processing device 10 that increases the user's motivation to learn a language by using a generated avatar as a conversation partner during language learning.

[0018] (Processing executed by information processing device 10) Here, the processing performed by the information processing device 10 will be described. Specifically, as shown in FIG. 1, the information processing device 10 generates an avatar according to the learning conditions of a user who is learning a language. Then, the information processing device 10 generates a trained model that generates a response sentence in response to an input of a conversation sentence acquired from the user, in response to the input of a conversation sentence acquired from the user. Then, the information processing device 10 outputs the generated response sentence from the avatar.

[0019] That is, the information processing device 10 takes in various data such as data about the user, such as the user's language level, data about a language model for pre-training a generation model that generates response sentences, and data about the situation in which the language learning will take place, and generates an avatar suitable for the user.The information processing device 10 then provides a comfortable language learning environment for the user by conversing with the user via the avatar suitable for the user.

[0020] Here, a language learning example provided by the information processing device 10 will be described with reference to Fig. 2. First, the information processing device 10 establishes a plug-in connection with the language information database 30 (step S1). For example, the information processing device 10 performs machine learning on a learned language model using data stored in the language information database 30 as training data, and generates a language model suitable for language learning.

[0021] In this state, the information processing device 10 accepts a selection of learning conditions for the user's language learning from the terminal device 20 (step S2). For example, the information processing device 10 accepts a selection that the user will study English.

[0022] Next, the information processing device 10 generates an avatar appropriately according to the user's selection of learning conditions such as the user's language level and language learning progress (step S3). Subsequently, the information processing device 10 generates a conversation response sentence based on the user's selection (step S4). Thereafter, the information processing device 10 outputs the generated response sentence from the avatar (step S5).

[0023] In this way, the information processing device 10 carries out a language learning conversation between the user and the avatar (step S6). For example, when the avatar outputs "Hello! How are you?" as voice data to the terminal device 20 as a generated response sentence, the user inputs voice data to the terminal device 20 in response to the avatar's question, saying "I'm fine. And you?". By repeating the above exchange, a conversation is carried out between the user and the avatar.

[0024] Finally, the information processing device 10 outputs feedback on the current language learning to the user (step S7). For example, the information processing device 10 generates, as feedback, video data of a video in which a conversation scene that took place during the language learning has been captured and text data in which the conversation has been transcribed, and outputs these to the terminal device 20. This provides an information processing device 10 that increases the user's motivation for language learning.

[0025] (1-2. Functional configuration of information processing device 10) Next, the functional configuration of the information processing device 10 will be described with reference to Fig. 3. Fig. 3 is an example of a functional block diagram showing the functional configuration of the information processing device 10 according to the first embodiment. As shown in Fig. 3, the information processing device 10 according to the first embodiment has a communication unit 11, a storage unit 12, and a control unit 13.

[0026] The communication unit 11 controls communication related to various information. For example, the communication unit 11 transmits and receives information to and from the language information database 30 and the like via the network N.

[0027] The storage unit 12 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk drive (HDD), a solid state drive (SSD), or an optical disk. The storage unit 12 stores various programs and various data. The storage unit 12 has a user information DB 12a, a language model 12b, and a situation DB 12c.

[0028] The user information DB 12a stores user information and data related to the language selected by the user that the user wishes to learn. FIG. 4 is a diagram illustrating the user information DB. As shown in FIG. 4, the user information DB 12a stores a "user name" and a "language" in association with each other. The "user name" stored here is information that identifies the user who is learning a language. The "language" is information that identifies the language that the user wishes to learn.

[0029] Using the example of FIG. 4, the user information DB 12a stores "User A, English," "User B, German," and "User C, Spanish" as "user name, language." That is, the user information DB 12a stores that "User A" selected "English" as the language he or she wants to learn. The user information DB 12a also stores that "User B" selected "German" as the language he or she wants to learn, and stores that "User C" selected "Spanish" as the language he or she wants to learn. In this way, the information processing device 10 can use the information stored in memory which language the user selected to help with language learning from the next time onward.

[0030] The language model 12b is a trained model that generates a response sentence in response to an input of a conversation sentence. Specifically, the language model 12b is generated by machine learning using conversation data from which a conversation is established as training data. For example, the language model 12b is generated by machine learning using an utterance sentence such as a question as an explanatory variable and a response sentence to the utterance sentence as a target variable. The language model 12b may be generated by the information processing device 10 by machine learning, or may be generated by another device. The information stored here may be trained parameters for constructing the trained language model 12b.

[0031] The situation DB 12c stores data related to situations that the user has selected and that the user wishes to learn. FIG. 5 is a diagram illustrating the situation DB. As shown in FIG. 5, the situation DB 12c stores a "user name," a "language," and a "situation" in association with each other. The "user name" stored here is information that identifies the user who is learning a language. The "language" is information that identifies the language that the user wishes to learn. The "situation" is information that identifies the situation that the user wishes to learn.

[0032] Using the example of FIG. 5, the situation DB 12c stores "User A, English, Cafe," "User B, German, Travel," and "User C, Spanish, Airport" as "User Name, Language, Situation." That is, the situation DB 12c stores that "User A" selected "Cafe" as the situation he wants to learn about in "English." The situation DB 12c also stores that "User B" selected "Travel" as the situation he wants to learn about in "German," and that "User C" selected "Airport" as the situation he wants to learn about in "Spanish."

[0033] 4, the control unit 13 of the information processing device 10 will now be described. The control unit 13 has an internal memory for storing programs that define various processing procedures and required data, and executes various processes using these. Here, the control unit 13 is, for example, an electronic circuit such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0034] The control unit 13 includes a receiving unit 13a, an acquiring unit 13b, a generating unit 13c, a specifying unit 13d, and an output control unit 13e.

[0035] The receiving unit 13a receives learning conditions including the language to be learned for language learning. After that, the information processing device 10 stores the information received by the receiving unit 13a in the user information DB 12a and the situation DB 12c.

[0036] Here, the screen on which the user selects the learning conditions will be described with reference to Fig. 6. Fig. 6 is an example of the display of the screen on which the user selects the learning conditions in the first embodiment. The screen in Fig. 6 is a screen that the reception unit 13a displays on the terminal device 20 used by the user.

[0037] As shown in Fig. 6, the reception unit 13a displays a screen including an area 41 for selecting a "situation" such as self-introduction or interview practice, and an area 42 for selecting a "language to learn" such as English. In the example of Fig. 6, the reception unit 13a receives a selection on this screen of "cafe" as the situation and "English" as the language the user wants to practice.

[0038] The acquisition unit 13b acquires data of conversational sentences uttered by the user. For example, the acquisition unit 13b acquires voice data or text data input by the user to the terminal device 20. Then, the acquisition unit 13b outputs the acquired data to the information processing device 10.

[0039] The generation unit 13c generates an avatar, generates responses to be spoken by the avatar, and generates text data of a series of conversations with the user. The processing of the generation unit 13c will be described later, but briefly, the generation unit 13c generates an avatar according to the learning conditions accepted by the acceptance unit 13a. The generation unit 13c also inputs audio data of conversations spoken by the user into the language model 12b and acquires the responses generated by the language model 12b. When language learning using the avatar is completed, the generation unit 13c also generates text data of the conversations spoken by the user and the responses spoken by the avatar in response to the conversations.

[0040] Returning to the explanation of FIG. 3, when language learning using an avatar is completed, the identification unit 13d identifies the user's learning level based on the conversational sentence spoken by the user and the response sentence spoken by the avatar in response to the conversational sentence. For example, the identification unit 13d identifies the user's learning level using the correctness or incorrectness of the user's response to a question, the number of words used by the user, the level of the words, etc. The identified learning level may be a five-point scale, a score, or a correspondence with another language test. The identification unit 13d stores the identified information in the user information DB 12a.

[0041] The output control unit 13e outputs the generated response sentence from an avatar. Specifically, the output control unit 13e can generate character data of the response sentence generated by the generation unit 13c and output it from an avatar, or can generate audio data of the response sentence generated by the generation unit 13c and output it from an avatar.

[0042] Here, the output control unit 13e can cause the avatar to perform an action corresponding to the situation selected by the user, while outputting a response sentence from the avatar. For example, when "cafe" is selected as the situation and a situation of ordering coffee occurs, the output control unit 13e outputs voice data along with an action of showing a menu. As another example, when "restaurant" is selected as the situation and a situation of paying the bill occurs, the output control unit 13e can output voice data along with an action of showing the amount or an action of receiving a card.

[0043] Furthermore, the output control unit 13e can output, as feedback for the user's language learning, text data generated by the generation unit 13c, including a conversational sentence spoken by the user and a response sentence spoken by the avatar in response to the conversational sentence. At this time, the output control unit 13e can also output video data of a conversation between the user and the avatar, captured while the language learning is being performed, together with the text data.

[0044] (Example) Here, a specific example of a conversation between a user and an avatar will be described with reference to Figs. 7 and 8. Fig. 7 is an example of the display of an avatar generated in the first embodiment. The screen in Fig. 7 is a screen that the generation unit 13c causes to be displayed on the terminal device 20 used by the user. The generation unit 13c displays the generated avatar in area 43 in Fig. 7. In addition, on the terminal device 20, area 44 for displaying a user who is learning a language is displayed on the screen together with area 43. In this way, the information processing device 10 generates an environment in which a conversation can take place between the user and the avatar.

[0045] Next, an example of a conversation between a user and an avatar will be described with reference to Fig. 8. Fig. 8 is a diagram illustrating an example of a conversation between a user and an avatar according to the first embodiment. Fig. 8 shows an example of an exchange between a user and an avatar generated according to the language and situation selected by the user. The following description will be given using an example of a conversation about ordering a drink at a coffee shop.

[0046] First, the generation unit 13c outputs voice data from an avatar saying "Welcome to GPT coffee shop! What can I get for you?" as voice indicating that language learning has started (step S51).

[0047] Next, the acquiring unit 13b acquires, from the terminal device 20 that output the voice data, the voice data "Well, Can I get a small iced Americano with cream please?" uttered by the user in response to the voice data of the avatar (step S52).

[0048] Then, the generation unit 13c inputs the acquired voice data "Well, Can I get a small iced Americano with cream please?" into the language model 12b to generate the next conversational sentence "Of course! That will be $3. Have a great coffee!", and the output control unit 13e causes the avatar to speak the generated conversational sentence "Of course! That will be $3. Have a great coffee!" (step S53).

[0049] Thereafter, the acquiring unit 13b acquires, from the terminal device 20, voice data of the user saying "Thank you!" as a response to the voice data of the avatar (step S54). As described above, conversation learning is performed between the user and the avatar in the situation specified by the user and in the language specified by the user.

[0050] Next, a specific example of feedback will be described with reference to Fig. 9. Fig. 9 shows an example of a display of a feedback screen for learning between a user and an avatar in the first embodiment. When a conversation between a user and an avatar ends, the generation unit 13c generates character data for each of the conversation sentence spoken by the user and the response sentence spoken by the avatar in response to the conversation sentence.

[0051] Then, the output control unit 13e outputs a feedback screen including an area 45 for displaying video data and an area 46 for displaying text data to the screen of the terminal device 20. For example, video data captured while the user is learning a language with the avatar is displayed in the area 45. Also, the area 45 displays text data that shows a series of conversations between the user and the avatar while they are learning a language, in which text data of conversational sentences uttered by the user and text data of response sentences given by the avatar are associated with each other.

[0052] As described above, by displaying such a feedback screen, the information processing device 10 can provide information that the user can use for review, which is expected to increase the user's motivation to learn.

[0053] (1-3. Processing Procedure of the First Embodiment) Next, an example of a processing procedure by the information processing device 10 according to the first embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing an example of a processing flow in the information processing device 10 according to the first embodiment.

[0054] 10, the information processing device 10 receives a selection of learning conditions, such as a language and a situation, from the user using the terminal device 20 (step S101). Next, the information processing device 10 generates an avatar based on the selection of learning conditions (step S102).

[0055] Next, the information processing device 10 causes the generated avatar to output a response sentence in a specific language (step S103). If the conversation between the user and the avatar has ended (step S104, Yes), the information processing device 10 outputs feedback on the conversation to the terminal device 20 and ends the process (step S105). On the other hand, if the conversation between the user and the avatar has not ended (step S104, No), the information processing device 10 continues the conversation between the user and the avatar. If the conversation has ended, the information processing device 10 outputs feedback on the conversation to the terminal device 20 (step S105) and ends the process.

[0056] (1-4. Effects of the First Embodiment) In the first embodiment, the information processing device 10 accepts a selection of a language that the user wants to learn, generates an avatar based on the accepted user selection, engages in a conversation with the user, and after the conversation ends, outputs feedback of the conversation to the terminal device 20. In this way, the information processing device 10 can improve the user's motivation to learn the language.

[0057] Furthermore, with the information processing device 10, the conversation partner in the language to be learned is not a human but an avatar, so that even if the user is not confident in speaking the language to be learned, the user can learn without feeling nervous.

[0058] Furthermore, if the information processing device 10 uses words or grammar of the language the user wants to learn incorrectly, the avatar will point this out as appropriate, eliminating the user's anxiety that the intended message may not be conveyed if the words or grammar are used incorrectly.

[0059] Furthermore, with the information processing device 10, the conversation partner in the language you want to learn is not a human but an avatar, so you can ask the same questions over and over again, which can alleviate your anxiety.

[0060] Furthermore, the information processing device 10 generates text data for the conversational sentence spoken by the user after the conversation ends and the response sentence spoken by the avatar in response to the conversational sentence, and outputs the data to the terminal device 20 in association with video data captured during the language learning process, thereby allowing the user to check the conversation they have practiced, providing an opportunity for the user to review.

[0061] Furthermore, the information processing device 10 can provide the user with a conversation with an avatar through the above-described processing, and speed up computer processing is realized.

[0062] <Second embodiment> (2-1. Overview of the Second Embodiment) In the first embodiment, an example of language learning suited to a situation such as a restaurant or cafe was described. However, the disclosed information processing device 10 can also order products directly from language learning by linking with an external server. Therefore, in the second embodiment, an example of placing an order with a store after a conversation between a user and an avatar is completed will be described. FIG. 11 is a diagram illustrating an example of processing according to the second embodiment. As shown in FIG. 11, the system configuration is the same as that shown in FIG. 1 described in the first embodiment, and therefore a detailed description thereof will be omitted.

[0063] (2-2. Processing Executed by Information Processing Device 10) Here, the processing performed by the information processing device 10 will be described. Specifically, as shown in FIG. 11, the information processing device 10 generates an avatar according to the learning conditions of a user who is learning a language. Next, the information processing device 10 generates a response sentence in response to input of a conversation sentence acquired from the user in a trained model that generates a response sentence in response to input of a conversation sentence acquired from the user. Next, the information processing device 10 outputs the generated response sentence from the avatar. Then, after the conversation between the user and the avatar ends, the information processing device 10 outputs to the terminal device 20 a screen that allows the user to select whether or not to order a product at an actual store.

[0064] That is, the information processing device 10 takes in various data including data related to the user's order and outputs a screen for selecting whether or not to transition to an ordering screen at a store. Then, by outputting an ordering screen at a store related to a conversation between the user and an avatar, the information processing device 10 connects language learning in an application or system with real life, providing the user with a successful experience.

[0065] (2-3. Configuration Functions of Information Processing Device 10) Next, the configuration functions of the information processing device 10 will be described with reference to Fig. 12. Fig. 12 is an example of a block diagram showing the functional configuration of the information processing device 10 according to the second embodiment. As shown in Fig. 12, the information processing device 10 according to the first embodiment has a communication unit 11, a storage unit 12, and a control unit 13.

[0066] 12, the information processing device 10 according to the second embodiment has functions equivalent to those of the communication unit 11, the user information DB 12a in the storage unit 12, the language model 12b, the situation DB 12c, the reception unit 13a, the acquisition unit 13b, the generation unit 13c, and the identification unit 13d in the control unit 13, which are described in FIG. 3, and therefore detailed description thereof will be omitted. Here, the order information DB 12d in the storage unit 12 and the output control unit 13e, which are different from those in the first embodiment, will be described.

[0067] The order information DB 12d stores data related to the order of a product selected by the user after the conversation ends. FIG. 13 is a diagram illustrating the order information DB. As shown in FIG. 13, the order information DB 12d stores a "user name," an "order," and a "store" in association with each other. The "user name" is information that identifies a user who is studying a language. The "order" is information that identifies a product order placed by the user. The "store" is information that identifies the store from which the product is ordered.

[0068] Using the example of FIG. 13, the order information DB 12d stores "User A, Yes, Store D," "User B, No, Store E," and "User C, No, Store F" as "user name, order, store." That is, as shown in FIG. 13, the order information DB 12d stores that "User A" placed an order "yes" and selected "Store D." The order information DB 12d also stores that "User B" placed an order "no" and did not select "Store E." The order information DB 12d also stores that "User C" placed an order "no" and did not select "Store F."

[0069] In addition to the processing described in the first embodiment, the output control unit 13e outputs a screen for selecting whether or not to place an order with the store after the conversation between the user and the avatar ends.

[0070] Here, a specific example of ordering a product will be described with reference to FIG. 14. FIG. 14 is an example showing an order screen according to the second embodiment. For example, if the content of the conversation between the user and the avatar is "ordering coffee in English at a cafe," the output control unit 13e outputs a screen for selecting whether or not to order coffee at an actual store. If the user selects "Yes" on the screen of FIG. 14, the information processing device 10 transitions to an order screen for store D, and if the user selects "No," learning ends. Thereafter, the information processing device 10 stores information related to the order in the order information DB 12d.

[0071] (2-4. Processing Procedure of the Second Embodiment) Next, an example of a processing procedure by the information processing device 10 according to the second embodiment will be described with reference to Fig. 15. Fig. 15 is a flowchart showing an example of a processing flow in the information processing device 10 according to the second embodiment.

[0072] 15, the information processing device 10 receives a selection of learning conditions, such as a language and a situation, from the user using the terminal device 20 (step S101). Next, the information processing device 10 generates an avatar based on the selection of learning conditions (step S102).

[0073] Next, the information processing device 10 causes the generated avatar to output a response sentence in a specific language (step S103). Then, when the conversation between the user and the avatar has ended (step S104, Yes), the information processing device 10 outputs a screen to the terminal device 20 that allows the user to select whether or not to order a product (step S104.1).

[0074] Then, if the user places an order (step S104.1, Yes), the information processing device 10 transitions to the store's order screen (step S104.2). On the other hand, if the user does not place an order (step S104.1, No), the information processing device 10 outputs learning feedback to the terminal device 20 (step S105) and ends the processing. Furthermore, when the order at the store is completed, the information processing device 10 outputs learning feedback to the terminal device 20 (step S105) and ends the processing.

[0075] (2-5. Effects of the Second Embodiment) For example, in the case of general language learning, there are learning methods that simulate ordering in a store, but users have few opportunities to actually order in English in a store, and may not have a successful experience. In response to this, the information processing device 10 accepts the user's selection of the language they want to learn, generates an avatar based on the user's selection, and engages in a conversation with the user. After the conversation ends, the terminal device 20 displays a selection screen asking the user whether or not to place an order in an actual store. If the order is completed or not placed, feedback on the conversation is output. Therefore, the information processing device 10 can connect language learning in an application or system with real life, providing the user with a successful experience.

[0076] <Third embodiment> (3-1. Overview of the Third Embodiment) In the above embodiments, an example has been described in which an avatar converses with a user using a trained language model 12b. However, the disclosed information processing device 10 can also tune the language model 12b to suit each user depending on the user's language ability, frequency of language learning, etc.

[0077] That is, the information processing device 10 can also re-learn (train) the avatar using the language learning history between the user and the avatar. Therefore, in the third embodiment, avatar re-learning will be described. Fig. 16 is a diagram illustrating an example of processing according to the third embodiment. As shown in Fig. 16, the system configuration is the same as Fig. 1 described in the first embodiment, and therefore detailed description will be omitted.

[0078] Here, the processing performed by the information processing device 10 will be described. Specifically, as shown in Fig. 16, the information processing device 10 executes the processes of generating an avatar, generating a response to a conversational sentence acquired from a user, and outputting the response sentence from the avatar, similar to the first embodiment. Thereafter, unlike the first embodiment, the information processing device 10 re-learns the language model 12b using the language learning history of the user and the avatar.

[0079] That is, the information processing device 10 re-learns the language model 12b using actual conversational sentences between a user and an avatar as training data. Therefore, the information processing device 10 can generate a language model 12b that is tuned to the habits and level of the user, and can provide language learning tailored to each user.

[0080] (3-2. Configuration Functions of Information Processing Device 10) The information processing device 10 according to the third embodiment has the same communication unit 11, storage unit 12, and control unit 13 as those described in Fig. 3, and therefore detailed description thereof will be omitted. Here, re-learning of a language model 12b, which is different from the above-described embodiment, will be described.

[0081] (3-3. Specific examples) Here, the relearning of the language model 12b will be described with reference to FIG. 17. FIG. 17 is an example of a diagram showing the relearning of the language model 12b according to the third embodiment. For example, the information processing device 10 stores in the user information DB 12a that the number of words used by the user during the first language study was "100," the level was "A1," and the number of turns of the user's conversation was "2." The information processing device 10 uses the information stored in the user information DB 12a to perform the relearning of the language model 12b.

[0082] Then, when performing the second language learning, the information processing device 10 uses the retrained language model 12b to generate an avatar with a higher level than the first one, and has a conversation with the user. For example, the information processing device 10 increases the number of words used by the avatar generated by the retrained language model 12b to "150," the level to "A2," and the number of conversation turns of the avatar to "4." This is expected to increase the number of words used in conversation with the user and the number of conversations.

[0083] (3-4. Effects of the Third Embodiment) In the third embodiment, the information processing device 10 receives a selection of a language learning program that a user wants to learn, generates an avatar based on the received user selection, and has a conversation with the user. The avatar then retrains using information about the conversation between the user and the avatar. This allows the information processing device 10 to provide language learning programs tailored to each user.

[0084] For example, in the case of general language learning, there are language learning methods that use applications or systems, but these methods may not provide language learning that is suitable for each user. In response to this, the information processing device 10 re-trains the language model 12b using user information after a conversation between the user and an avatar ends, and generates an avatar with a higher level for the next language learning session and outputs it to the terminal device 20. Therefore, the information processing device 10 provides language learning that is suitable for each user by re-training the language model 12b using information on the user's past language learning.

[0085] <Fourth embodiment> (4-1. Overview of the Fourth Embodiment) In the above-described embodiments, examples of language learning in situations such as restaurants and cafes have been described. However, the disclosed information processing device 10 can also be used for purposes other than language learning, such as interview practice. Therefore, in the fourth embodiment, interview practice will be described. The system configuration of the fourth embodiment is the same as that of FIG. 1 described in the first embodiment, and therefore a detailed description thereof will be omitted.

[0086] (4-2. Functional configuration of information processing device 10) Next, the configuration functions of the information processing device 10 will be described. The information processing device 10 according to the fourth embodiment has functions equivalent to those of the communication unit 11, the situation DB 12c of the storage unit 12, the acquisition unit 13b, the generation unit 13c, and the identification unit 13d of the control unit 13, which are described in Fig. 3, and therefore detailed description thereof will be omitted. Here, the user information DB 12a, the language model 12b, the situation DB 12c of the storage unit 12, the reception unit 13a, and the output control unit 13e, which are different from those of the first embodiment, will be described.

[0087] The user information DB 12a stores user information and data related to the language selected by the user that the user wishes to learn. FIG. 18 is a diagram illustrating the user information DB. As shown in FIG. 18, the user information DB 12a stores a "user name" and a "language" in association with each other. The "user name" stored here is information that identifies the user who is learning a language. The "language" is information that identifies the language that the user wishes to learn.

[0088] Using the example of FIG. 18 as an explanation, the user information DB 12a stores "User G, Japanese," "User H, Japanese," and "User I, English" as "user name, language." That is, the user information DB 12a stores that "User G" selected "Japanese" as the language he wants to learn. The user information DB 12a also stores that "User H" selected "Japanese" as the language he wants to learn, and stores that "User I" selected "English" as the language he wants to learn. In this way, the information processing device 10 can use the information selected by the user to learn a language from the next time onward by storing the language selected by the user.

[0089] The situation DB 12c stores data related to situations that the user has selected and that the user wishes to learn. FIG. 19 is a diagram illustrating the situation DB. As shown in FIG. 19, the situation DB 12c stores a "user name," a "language," and a "situation" in association with each other. The "user name" stored here is information that identifies the user who is learning a language. The "language" is information that identifies the language that the user wishes to learn. The "situation" is information that identifies the situation that the user wishes to practice.

[0090] Explaining the example of FIG. 19, the situation DB 12c stores "User G, Japanese, job hunting interview practice," "User H, Japanese, qualification exam interview practice," and "User I, English, presentation practice" as "user name, language, situation." That is, the situation DB 12c stores that "User G" selected "job hunting interview practice" as the situation for which he wants to practice in "Japanese." The situation DB 12c also stores that "User H" selected "qualification exam interview practice" as the situation for which he wants to practice in "Japanese," and that "User I" selected "presentation practice" as the situation for which he wants to practice in "English."

[0091] (Example) Next, a screen on which a user selects learning conditions will be described with reference to FIG. 20. FIG. 20 is an example of a display of a screen selected by a user according to the fourth embodiment. The screen of FIG. 20 is a screen that the reception unit 13a displays on the terminal device 20 used by the user. The reception unit 13a displays a screen including an area for selecting a "situation" such as self-introduction or interview practice, and an area for selecting a "language to learn" such as Japanese. The reception unit 13a then receives a selection of "interview practice" in area 47 of FIG. 20 as the situation on the screen, and "Japanese" in area 48 of FIG. 20 as the language the user wants to learn.

[0092] The information processing device 10 outputs the avatar generated by the generation unit 13c based on the information received by the reception unit 13a to the terminal device 20 used by the user. The information processing device 10 then conducts interview practice between the user and the avatar. In this way, the information processing device 10 provides the user with interview practice as a use other than language learning.

[0093] For example, as a conversation for an interview practice, first, the generation unit 13c outputs voice data from the avatar saying, "The interview will now begin. Thank you for your cooperation." as voice data indicating that the interview practice has started.

[0094] Next, the acquiring unit 13b acquires, from the terminal device 20 that has output the voice data, the voice data of "Yes, thank you very much," spoken by the user in response to the voice data of the avatar.

[0095] Then, the generation unit 13c inputs the acquired voice data "Yes, thank you very much" into the language model 12b to generate the next conversational sentence "Please introduce yourself.", and the output control unit 13e causes the avatar to speak the generated conversational sentence "Please introduce yourself."

[0096] Thereafter, the acquiring unit 13b acquires from the terminal device 20, as a response to the avatar's voice data, the user's voice data saying, "My name is G. I belong to the K department of J University."

[0097] Then, the generation unit 13c inputs the acquired voice data "My name is G and I belong to the K department of J university" into the language model 12b to generate the next conversational sentence "This concludes the interview practice.", and the output control unit 13e causes the avatar to speak the generated conversational sentence "This concludes the interview practice." As described above, the interview practice is carried out between the user and the avatar in a situation specified by the user and in a language specified by the user.

[0098] (4-3. Effects of the Fourth Embodiment) In the fourth embodiment, the information processing device 10 receives a selection of an interview practice that the user wants to learn, generates an avatar based on the received user selection, and has a conversation with the user. This allows the information processing device 10 to easily provide the user with interview practice.

[0099] Furthermore, when comparing general interview practice with the information processing device 10, applications and systems that provide general interview practice have difficulty providing interview practice that is suitable for each user. For example, if a user wants to practice for a job interview, there are two methods: face-to-face practice and practice using an application or system. However, the user must take the time to find someone to practice with. In addition, the user may feel nervous or anxious about practicing for an interview. Therefore, the information processing device 10 can provide interview practice that is suitable for each user by using an avatar to practice for the interview.

[0100] Furthermore, the information processing device 10 generates an avatar for interview practice on the terminal used by the user, thereby reducing the time required to find a partner for interview practice.

[0101] Furthermore, with the information processing device 10, the person who is to practice interviewing is not a human being but an avatar, so that the person can practice interviewing without feeling nervous or anxious.

[0102] <Fifth embodiment> Although the embodiments of the present invention have been described above, the present invention may be embodied in various different forms other than the above-described embodiments.

[0103] (Usage form) The output control unit 13e of the information processing device 10 described in the first embodiment outputs a response sentence from an avatar to the user who is studying a language again, using at least one of the avatar's appearance, speaking speed, and colloquialized version of the response sentence according to the learning level identified at the end of the previous language study.

[0104] For example, when the user's language learning level increases, the information processing device 10 may change the avatar's appearance, including clothing, to resemble that of a local person, or may change the speaking speed of the language learner to be closer to that of a local person, or may use abbreviations and dialects in the output response sentences, etc., as an avatar according to the learning level.

[0105] Furthermore, when a user makes a vocabulary or grammatical error while conversing with an avatar, the information processing device 10 described in the first embodiment corrects the error and outputs the error from the avatar to the terminal device 20. For example, when there is a vocabulary error, the information processing device 10 outputs a notice of the error as voice data or text data to the terminal device 20. Furthermore, when there is an error in the user's response to the avatar, the information processing device 10 outputs the notice of the error as voice data or text data to the terminal device 20, and the avatar repeatedly points out the error until the user corrects the error and the response is corrected.

[0106] (Hardware configuration) The information processing device 10, the terminal device 20, and the language information database 30 included in the route determination system 1 according to the embodiment may be realized by a computer 1000 configured as shown in FIG. 21, for example. The determination device 200 will be described below as an example. FIG. 21 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the information processing device 10. The computer 1000 may have a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.

[0107] The CPU 1100 may operate and control each unit based on a program stored in the ROM 1300 or the HDD 1400. The ROM 1300 may store a boot program executed by the CPU 1100 when the computer 1000 starts up, a program dependent on the hardware of the computer 1000, and the like.

[0108] The HDD 1400 may store programs executed by the CPU 1100, data used by such programs, etc. The communication interface 1500 may receive data from other devices via the communication network 50 and transmit the data to the CPU 1100. The communication interface 1500 may transmit data generated by the CPU 1100 to other devices via the communication network 50.

[0109] The CPU 1100 may control output devices such as a display and a printer, and input devices such as a keyboard and a mouse, via the input / output interface 1600. The CPU 1100 may acquire data from the input devices via the input / output interface 1600. The CPU 1100 may also output generated data to the output devices via the input / output interface 1600.

[0110] Media interface 1700 may read a program or data stored in recording medium 1800 and provide it to CPU 1100 via RAM 1200. CPU 1100 may load the program or data from recording medium 1800 onto RAM 1200 via media interface 1700 and execute the loaded program. Recording medium 1800 may be, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0111] For example, when the computer 1000 functions as the information processing device 10 according to the embodiment, the CPU 1100 of the computer 1000 may implement the functions of the control unit 230 by executing a program loaded onto the RAM 1200. Furthermore, the HDD 1400 may store data in the storage unit 120. The CPU 1100 may read and execute these programs from the recording medium 1800. The CPU 1100 may also obtain these programs from another device via the communication network 50.

[0112] (others) Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0113] For example, when the above embodiment includes a plurality of information processing devices 10, the plurality of terminal devices 20 may be different devices. In other words, as long as the functions of the terminal device can be realized, the plurality of terminal devices 20 do not have to be the same device. For example, the shape and functions of the device may differ depending on the situation in which the terminal device 20 is installed or mounted.

[0114] The above describes in detail the embodiments of the present application based on several drawings, but these are merely examples, and the present invention can be implemented in other forms that include the embodiments described in the Disclosure of the Invention section and that have been modified and improved in various ways based on the knowledge of those skilled in the art.

[0115] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, a decision unit can be read as a decision means or a decision circuit, etc. [Explanation of symbols]

[0116] 10. Information processing equipment 20 Terminal equipment 30 Language Information Database 11 Communications Department 12 Storage section 12a User Information DB 12b Language Model 12c Situation DB 12d Order Information DB 13 Control Unit 13a Reception 13b Acquisition part 13c generator 13d Specific part 13e Output control section N Network

Claims

1. a first generation unit that generates an avatar according to a learning condition of a user who is learning a language; a second generation unit that generates a response sentence by inputting a conversation sentence acquired from the user into a trained model that generates a response sentence in response to an input of a conversation sentence; an output control unit that outputs the generated response sentence from the avatar; An information processing device comprising:

2. an acquisition unit that acquires voice data of a conversation sentence uttered by the user; The second generation unit inputting the speech data into the trained model to generate the response sentence; The output control unit causing the avatar to speak the response sentence; The information processing device according to claim 1 .

3. The learning method further includes a receiving unit that receives the learning conditions including a language to be learned for the language learning. The acquisition unit acquiring the speech data spoken in the language to be learned; The second generation unit inputting the speech data into the trained model to generate the response sentence; The output control unit causing the avatar to speak the response sentence in the language to be learned; The information processing device according to claim 2 .

4. a receiving unit that receives the learning conditions including a situation in which the language learning is to be performed; The output control unit causing the avatar to perform an action according to the situation and uttering the response sentence; The information processing device according to claim 2 .

5. The output control unit generating character data of a conversation sentence spoken by the user and a response sentence spoken by the avatar in response to the conversation sentence when the language learning using the avatar is completed; The character data of the conversation sentence and the character data of the response sentence are output in association with each other. The information processing device according to claim 1 .

6. The output control unit and outputting video data in which the user speaks and the avatar responds, the video data being captured while the language learning is being performed, in association with the character data of the conversation sentence and the character data of the response sentence. The information processing device according to claim 5 .

7. a determination unit that, when the language learning using the avatar is completed, determines a learning level of the user based on a conversation sentence spoken by the user and a response sentence spoken by the avatar in response to the conversation sentence; The output control unit outputting the response sentence from the avatar to the user who is studying the language again, using at least one of the appearance of the avatar, the speech speed, and a colloquialized version of the response sentence, according to the study level identified at the end of the previous language study; The information processing device according to claim 2 .

8. a first generation step of generating an avatar according to a learning condition of a user who is learning a language; a second generation step of inputting a conversational sentence acquired from the user into a trained model that generates a response sentence in response to the input of a conversational sentence, and generating a response sentence; an output control step of outputting the generated response sentence from the avatar; An information processing method comprising:

9. a first generation step of generating an avatar according to a learning condition of a user who is learning a language; a second generation step of inputting a conversational sentence acquired from the user into a trained model that generates a response sentence in response to an input of a conversational sentence, and generating a response sentence; an output control step of outputting the generated response sentence from the avatar; An information processing program comprising:

Citation Information

Patent Citations

  • Foreign language conversation training system using computer

    JP2012215645A

  • System and method for learning foreign language

    KR1020110000307A

  • Circuit input solver, circuit input design method and medium storing thereof

    KR1020240072062A

  • C.i.c.

    KR102640880B1

  • System for Communication Skills Training Using Juxtaposition of Recorded Takes

    US20220028369A1