Foreign language learning support system and foreign language learning support method
The foreign language acquisition support system addresses the lack of non-verbal communication skill development in conventional methods by evaluating and providing feedback on facial expressions, gestures, and cultural context, enhancing overall communication abilities.
Patent Information
- Application Number
- JP2024104553
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-01-16
AI Technical Summary
Conventional shadowing devices primarily focus on improving language skills without considering non-verbal communication skills, which are crucial for effective communication.
A foreign language acquisition support system that includes imaging, evaluation, and display means to assess and improve non-verbal communication skills through facial expressions, gestures, and cultural background awareness, using machine learning to generate evaluation models and provide personalized learning materials.
Enhances users' non-verbal communication abilities, enabling more effective language learning by integrating non-verbal skill evaluation and feedback into the learning process.
Smart Images

Figure 2026005914000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a foreign language learning support system, a foreign language learning support method, and a program for causing a computer to execute the method, which support a user in learning a foreign language. [Background technology]
[0002] "Shadowing" is an effective method for learning a foreign language. "Shadowing" is a training method, primarily for simultaneous interpretation, in which you listen to a foreign language (such as English) and imitate its pronunciation. Unlike "repeating," which involves repeating the foreign language after listening, shadowing involves pronouncing the foreign language immediately after listening. The name comes from the fact that you follow the foreign language like a shadow. An example of a shadowing device for performing shadowing is described in Japanese Patent No. 5756555. This shadowing device records the learner's voice while reading aloud or shadowing, and then objectively evaluates the recorded voice. This is said to support independent learning and reduce the amount of work required by instructors. [Prior art documents] [Patent documents]
[0003] Patent No. 5756555 Summary of the Invention [Problem to be solved by the invention]
[0004] The ultimate goal of learning a foreign language is not simply to acquire the ability to listen and speak it, but to communicate effectively with others through language. In general, it is said that non-verbal elements (non-verbal communication skills) play a greater role in communication than language itself. Non-verbal elements are means of expression other than language, such as facial expressions, gestures, and body actions. These non-verbal elements not only complement and emphasize verbal messages, but also help us understand the other person's feelings and intentions. Furthermore, non-verbal elements are also important to avoid misunderstandings, as the same words can be interpreted in completely different ways depending on facial expressions, body language, and hand gestures.
[0005] Thus, in order to communicate effectively with others, it is important to acquire not only language skills but also non-verbal communication skills. Although there is no universal language, non-verbal communication skills are universal (for example, this is why Chaplin's silent films were accepted worldwide). However, the above-mentioned conventional shadowing devices are aimed only at improving language skills and are completely indifferent to non-verbal communication skills. The present invention was made in consideration of the problems with conventional shadowing devices, and aims to provide a foreign language acquisition support system, a foreign language acquisition support method, and a program for said method that enable users to improve their non-verbal communication skills when learning a foreign language, regardless of the foreign language learning method. [Means for solving the problem]
[0006] In order to achieve this object, the present invention provides, as a first aspect, a foreign language acquisition support system (500) for supporting a user in acquiring a foreign language, the foreign language acquisition support system (500) comprising: an imaging means (145) for imaging the user enunciating in accordance with the audio of a video; a first storage means (132D) for storing exemplary non-verbal communication skills of interlocutors during a conversation; a second storage means (132E) for storing a learned evaluation model for evaluating the interlocutors' non-verbal communication skills during a conversation; an evaluation means (131D) for comparing the exemplary non-verbal communication skills stored in the first storage means (132D) with the non-verbal communication skills of the user imaged by the imaging means (145) using the learned evaluation model (210B) stored in the second storage means (132E), and evaluating the non-verbal communication skills of the user; and a display means (131E, 142) for displaying the evaluation created by the evaluation means (131D). It is preferable that the evaluation means (131D) evaluates the non-verbal communication skills of the user based on at least one evaluation item of facial expression, gaze, gesture, body language, proxemics, physical appearance, visual focus, auditory information, and cultural background.
[0007] It is preferable that the evaluation means (131D) extracts features quantitatively indicating the characteristics of each of the evaluation items from the data of the image of the user captured by the imaging means (145), and compares these features with exemplary non-verbal communication skills stored in the first storage means (132D). It is preferable that the evaluation means (131D) uses the content of the user's conversation as language data in an auxiliary manner when evaluating the user's non-verbal communication skills. It is preferable that the foreign language acquisition assistance system (500) further comprises a cultural background information database storing cultural background information of each country or region, and an adjustment means for reading out the cultural background information of the user from the cultural background information database and adjusting the evaluation of the user's non-verbal communication skills by the evaluation means (131D) taking into account the cultural background information of the user.
[0008] The trained evaluation model (210B) is generated by machine learning to evaluate the non-verbal communication skills of a conversation partner using exemplary non-verbal communication skills stored in the first storage means (132D) as a criterion for judgment, and the input (220) to the trained evaluation model (210B) is image data of the user's non-verbal communication skills during conversation captured by the imaging means (145), and the output (230) from the trained evaluation model (210B) is preferably an evaluation of the user's non-verbal communication skills during conversation using the exemplary non-verbal communication skills as a criterion for judgment, by executing a predetermined calculation process. It is preferable that the foreign language acquisition support system (500) further comprises a curriculum creation means for creating a curriculum specialized for the user in order to compensate for the non-verbal communication skills of the user that have been poorly evaluated by the evaluation means (131D). It is preferable that the foreign language acquisition assistance system (500) further comprises a learning material creation means for creating learning materials in accordance with the curriculum created by the curriculum creation means.
[0009] For example, the learning material is a video, and the video is preferably an interactive video in which the user (320) and the characters (310) have a dialogue. The foreign language acquisition assistance system (500) further includes a subtitle generation means for displaying subtitles (330) on the screen of the video, and it is preferable that the subtitle generation means displays, when a non-verbal communication skill of the user (320) that has been poorly evaluated by the evaluation means (131D) appears during the dialogue, at least a subtitle (330) indicating that and advice regarding the non-verbal communication skill on the screen. The foreign language learning assistance system (500) further includes a subtitle audio conversion means for converting the subtitles (330) into audio, and the subtitle audio conversion means preferably outputs an audio version of the subtitles (330) from the screen instead of the subtitles (330) or together with the subtitles (330).
[0010] In a second aspect, the present invention provides a portable wireless communication device (100) incorporating the above foreign language learning assistance system (500). In a third aspect, the present invention provides a foreign language acquisition assistance method for assisting a user in acquiring a foreign language, comprising: a first step (S140) of imaging the user enunciating in accordance with the audio of a video; a second step (S210) of evaluating the user's nonverbal communication skills by comparing the nonverbal communication skills of the user imaged in the first step (S140) with exemplary nonverbal communication skills stored in a storage means (132D) that pre-stores exemplary facial expressions, gestures, and other exemplary nonverbal communication skills of interlocutors during conversation using a learned evaluation model (210B) that evaluates the interlocutors' nonverbal communication skills; and a third step (S220) of displaying the evaluation created in the second step (S210).
[0011] In the second step (S210), it is preferable that the user's non-verbal communication skills are evaluated based on at least one evaluation item of facial expressions, gaze, gestures, body language, proxemics, physical appearance, visual focus, auditory information, and cultural background. In the second step (S210), it is preferable that features quantitatively indicating the characteristics of each of the evaluation items are extracted from the image data of the user, and these features are compared with exemplary non-verbal communication skills stored in the storage means (132D). In the second step (S210), it is preferable that the content of the user's conversation is used as auxiliary language data when evaluating the user's non-verbal communication skills. In the second step (S210), it is preferable that the cultural background information of the user is read from a cultural background information database that stores cultural background information of each country or region, and the evaluation of the user's non-verbal communication skills is adjusted taking into account the cultural background information of the user.
[0012] The trained evaluation model (210B) is generated by machine learning to evaluate the non-verbal communication skills of the interlocutor using exemplary non-verbal communication skills stored in the storage means (132D) as a criterion for judgment, and the input (220) to the trained evaluation model (210B) is image data capturing the non-verbal communication skills of the user during conversation, and the output (230) from the trained evaluation model (210B) is preferably an evaluation of the non-verbal communication skills of the user during conversation using the exemplary non-verbal communication skills as a criterion for judgment by executing a predetermined calculation process. It is preferable that the foreign language acquisition support method further comprises a fourth step of creating a curriculum specialized for the user to compensate for the non-verbal communication skills of the user that were evaluated poorly in the second step (S210).
[0013] It is preferable that the foreign language acquisition assistance method further comprises a fifth step of creating learning materials in accordance with the curriculum created in the fourth step. For example, the learning material is a video, and the video is preferably an interactive video in which the user (320) and the characters (310) have a dialogue. It is preferable that the foreign language acquisition assistance method further comprises a fifth step of, if the non-verbal communication skill of the user that was poorly evaluated in the second step (S210) appears during the dialogue, displaying on the screen at least a subtitle (330) indicating that and advice regarding the non-verbal communication skill. It is preferable that the foreign language acquisition assistance method further comprises a sixth step of converting the subtitles (330) into audio, and a seventh step of playing the audio of the subtitles (330) from the screen instead of or together with the subtitles (330).
[0014] In a fourth aspect, the present invention provides a program for causing a computer to execute the above foreign language learning assistance method. In a fifth aspect, the present invention provides a portable wireless communication device (100) incorporating the above program. The reference numerals in parentheses are used to indicate correspondence with the embodiments described later, and are not intended to limit the scope of the claims. [Effects of the Invention]
[0015] In order to communicate effectively with others, in addition to verbal skills, non-verbal communication skills (facial expressions, gestures, hand movements, etc.) are also important elements. However, conventional systems (such as foreign language acquisition assistance devices, foreign language learning materials, and foreign language schools) have been completely indifferent to non-verbal communication skills. The foreign language learning assistance system according to the present invention enables users to improve their non-verbal communication skills. Specifically, it enables users to learn exemplary non-verbal communication skills during conversations, thereby improving the users' communication abilities. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a block diagram showing an example of the structure of a portable wireless communication device incorporating a foreign language learning assistance system according to a first embodiment of the present invention. [Figure 2] FIG. 10 is a conceptual diagram showing the configuration of a third program stored in an external memory. [Figure 3] FIG. 10 is a conceptual diagram showing functions of a third program stored in an external memory. [Figure 4] 2 is a flowchart showing the operation of the foreign language learning assistance system shown in FIG. [Figure 5] FIG. 10 is a conceptual diagram of a foreign language learning assistance system according to a fifth embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0017] (First embodiment) FIG. 1 is a block diagram showing an example of the structure of a portable wireless communication device 100 incorporating a foreign language learning assistance system 500 according to a first embodiment of the present invention. The portable wireless communication device 100 comprises a communication unit 110, a control unit 120, an external memory (hard disk) 130, an input / output unit 140, and an antenna 150, and of these, the control unit 120, the external memory 130, and the input / output unit 140 constitute a foreign language acquisition support system 500. The portable wireless communication device 100 is configured as a portable telephone device such as a smartphone, for example. The communication unit 110 is connected to an antenna 150, and transmits and receives data via the antenna 150 to and from other wireless communication devices via wireless communication. The communication unit 110 includes a wireless receiving unit 111 , a wireless transmitting unit 112 , and a changeover switch 113 .
[0018] Wireless receiving unit 111 demodulates data received from other wireless communication devices and sends the data to control unit 120. Wireless transmitting unit 112 modulates data output from control unit 120 and transmits the data to other mobile phone devices or other wireless communication devices via antenna 150. Changeover switch 113 receives a signal from control unit 120 and switches between transmission and reception in response to the signal. The control unit 120 is composed of a central processing unit (CPU) 121, a first memory 122 consisting of ROM, a second memory 123 consisting of RAM, an input interface 124 for transferring various commands and data input to the control unit 120 to the central processing unit 121, an output interface 125 for outputting the processing results executed by the central processing unit 121 to the outside, and a bus 126 connecting the central processing unit 121 to each of the first memory 122, the second memory 123, the input interface 124 and the output interface 125.
[0019] The first memory 122 stores various control programs executed by the central processing unit 121 and other non-rewritable data. The second memory 123 stores various data and parameters and provides a working area for the central processing unit 121. In other words, data or programs temporarily required for the central processing unit 121 to execute various control programs are read from the external memory 130 and temporarily stored in the second memory 123. The central processing unit 121 cooperates with an OS (Operating System) to control the overall operation of the portable wireless communication device 100. Specifically, it reads programs necessary for the operation of the portable wireless communication device 100 from the external memory 130 and executes these programs. In other words, the central processing unit 121 operates in accordance with the programs stored in the external memory 130. As will be described later, the central processing unit 121 outputs a predetermined result in response to a predetermined input via a trained evaluation model.
[0020] The input / output unit 140 is made up of an operation unit 141, a display 142, a speaker 143, a microphone (sound collector) 144, and a camera 145 as an imaging device. The operation unit 141 is made up of, for example, a numeric keypad, and various data are input to the control unit 120 via the operation unit 141 . The display 142 is, for example, a liquid crystal display, and displays the results of calculations performed by the control unit 120 (such as evaluation results, which will be described later) and other data on the screen. The voice (synthesized voice) generated by the control unit 120 is output via a speaker 143. The audio data collected by the microphone 144 and the image data captured by the camera 145 are sent to the control unit 120 .
[0021] The external memory (hard disk) 130 is composed of an application storage section 131 and a data storage section 132 . The data storage unit 132 is composed of a video storage unit 132A in which various videos collected to date are stored, a video audio storage unit 132B in which the audio of a video selected by the user is recorded as audio data, a user audio recording unit 132C in which the audio uttered by the user following the audio of the video is recorded as audio data, a first data storage unit 132D in which data on exemplary language skills (accurate pronunciation, accent, exemplary fluency, etc.) is stored and the exemplary non-verbal communication skills of the interlocutor during a conversation, a second data storage unit 132E in which a trained first evaluation model that evaluates the user's language skills and a trained second evaluation model that evaluates the interlocutor's non-verbal communication skills during a conversation, a user image data storage unit 132F in which an image of the user (specifically, the user's non-verbal communication skills) captured by a camera 145 is recorded as image data, and a sub-data storage unit 132G in which various data other than these video data are stored.
[0022] In this specification, the "non-verbal communication skills" of a conversational participant include at least the following items: (1) Facial expressions (targeting changes in emotions and subtle facial movements) (2) Gaze (targeting gaze direction, duration, etc.) (3) Gestures (hand and arm movements, etc.) (4) Body language (posture and body orientation) (5) Proxemics (targeting interpersonal distance and interpersonal angle) (6) Physical appearance (including clothing, hairstyle, accessories, etc.) The application unit 131 stores an OS (Operating System) 131S that controls the overall operation of the portable wireless communication device 100, a first program 131A, a second program 131B, a third program 131C, a fourth program 131D, and a fifth program 131E.
[0023] First program 131A constitutes a moving image audio recording means for recording the audio of a moving image selected by the user from the moving images stored in moving image storage section 132A in moving image audio storage section 132B. The second program 131B constitutes a user voice recording means that records the voice uttered by the user in accordance with the video selected by the user in the user voice storage unit 132C, and also constitutes a user image data recording means that records the user's non-verbal communication skills captured by the camera 145 in the user image data storage unit 132F. The third program 131C evaluates the user's language skills by comparing the exemplary language skills stored in the first data storage unit 132D with the user's voice data stored in the user voice storage unit 132C using the trained first evaluation model stored in the second data storage unit 132E, and creates an evaluation result.
[0024] The fourth program 131D constitutes an evaluation means that uses the trained second evaluation model stored in the second data storage unit 132E to compare exemplary nonverbal communication skills during conversation stored in the first data storage unit 132D with the user's nonverbal communication skills during conversation in image data (image data capturing the user's nonverbal communication skills during conversation) captured by the camera 145 and sent to the control unit 120, thereby creating an evaluation result that evaluates the user's nonverbal communication skills during conversation. In other words, the fourth program 131D evaluates the user's non-verbal communication skills during shadowing (or reading aloud, speaking, etc.).
[0025] The fifth program 131E has a function of displaying the evaluation results created by the third program 131C and the fourth program 131D on the display 142. The fifth program 131E and the display 142 constitute a display means. FIG. 2 is a conceptual diagram showing the configuration of the third program 131C. As shown in FIG. 2, the third program 131C is made up of a teacher data input program 1311, an evaluation model generation program 1312, and an evaluation result output program 1313. The teacher data input program 1311 is a program for inputting teacher data to the control unit 120 for generating an evaluation model by machine learning. The evaluation model generation program 1312 generates each evaluation model by machine learning using the training data input by the training data input program 1311, and outputs the generated trained evaluation models to the central processing unit 121. The evaluation result output program 1313 outputs the results of evaluation using the trained evaluation model generated by the evaluation model generation program 1312.
[0026] The fourth program 131D has the same configuration as the third program 131C. FIG. 3 is a conceptual diagram showing the function of the third program 131C. As shown in FIG. 3, training data 200 is sent to an evaluation model generation program 1312 by a training data input program 1311, and the evaluation model generation program 1312 generates an evaluation model 210 based on the training data 200. This evaluation model 210 continues to undergo machine learning using subsequent inputs as training data 200, and becomes a trained evaluation model 210. The first evaluation model 210A of the evaluation models 210 is generated by machine learning to evaluate whether the user's language skills are appropriate using exemplary language skills (accurate pronunciation and accent, exemplary fluency, appropriate vocabulary, etc.) stored in the first data storage unit 132D as criteria for judgment. The trained first evaluation model 210A receives input of the user's voice during conversation as voice data (input 220).
[0027] The trained first evaluation model 210A performs predetermined calculation processing to output 230 an evaluation of whether the user's language skills are appropriate, using exemplary language skills as a criterion. The evaluation result (output 230) generated by the trained first evaluation model 210A is sent to the control unit 120 by the evaluation result output program 1313, and then displayed on the display 142 by the fifth program 131E. The trained second evaluation model 210B is generated by machine learning to evaluate whether the user's non-verbal communication skills during a conversation are appropriate, using exemplary non-verbal communication skills during the conversation stored in the first data storage unit 132D as a criterion for judgment. Image data of the user's non-verbal communication skills captured by the camera 145 during conversation is introduced as input 220 to the trained second evaluation model 210B.
[0028] The trained second evaluation model 210B performs predetermined calculations to output 230 an evaluation of whether the user's non-verbal communication skills during conversation are appropriate, using as a criterion non-verbal communication skills during conversation that are generally considered exemplary. The evaluation result (output 230) generated by the trained second evaluation model 210B is sent to the control unit 120 by the evaluation result output program 1313, and then displayed on the display 142 by the fifth program 131E. 4 is a flowchart showing the operation of the portable radio communication device 100. The operation of the portable radio communication device 100 will be described below with reference to FIG. First, the user selects a moving image that he or she wishes to study (for example, shadow) from among the many moving images stored in moving image storage unit 132A (step S100).
[0029] Various videos are stored in video storage unit 132A. For example, users can select a video that suits their needs from videos categorized by genre (business scenes, shopping scenes, airplane scenes, etc.), language (major languages such as English and French as well as less commonly used minor languages), region (for example, the same English word may be pronounced differently in the UK, the US, and Australia), length (from videos over an hour to just a few minutes), etc. Next, before starting the study, the user specifies the evaluation level for his / her own language skills and non-verbal communication skills during conversation (step S110). For example, there are several evaluation levels, ranging from "lenient evaluation" to "strict evaluation," and users can specify the evaluation level that suits their own learning progress.
[0030] Thereafter, the user activates the sound recording device (not shown) and the camera 145 to start recording and sound recording (step S120). Next, the user begins learning. Specifically, the audio of the video is emitted from speaker 143, and the user pronounces the audio by imitating the audio of the video immediately after the audio of the video or at predetermined intervals. The audio emitted by the user is collected by microphone 144 and recorded in user audio storage unit 132C of data storage unit 132 via second program 131B (step S130). Along with recording the user's voice, the non-verbal communication skills of the user being studied are photographed by camera 145 (step S140) and recorded in user image data storage section 132F of data storage section 132 via second program 131B. The voice data of the user's voice and the image data of the user's non-verbal communication skills are sent to the control unit 120. At the same time as recording the user's voice and the user's non-verbal communication skills, the audio of the video is also recorded in the video audio storage unit 132B via the first program 131A (step S150).
[0031] The user's voice recorded in the user voice storage unit 132C is converted into characters (text) by a program (not shown) having a voice recognition function (step S160). When the learning is completed (step S170), the central processing unit 121 starts the third program 131C, and using the learned first evaluation model stored in the second data storage unit 132E, compares the exemplary language skills stored in the first data storage unit 132D with the user's actual voice recorded in the user voice storage unit 132C, and evaluates the user's pronunciation (language skills) (step S180).
[0032] Language skills will be assessed on at least the following points: (A) Pronunciation evaluation Specifically, the clarity and accuracy of each sound's pronunciation, the pronunciation of syllables, the distinction between vowels and consonants, and the placement of accents are evaluated. (B) Intonation evaluation Specifically, the intonation of sentences and phrases, changes in pitch, rhythm, appropriateness of stress and emphasis, and expression of emotion are evaluated. (C) Fluency assessment Specifically, the fluency, naturalness, coherence, pace, and smoothness of word transitions are evaluated. (D) Pointing out errors It detects slip-ups, mispronunciations, grammatical errors, poor vocabulary choices, and more. Thereafter, the central processing unit 121 displays the above evaluation results on the display 142 (step S190).
[0033] Table 1 is an example of the evaluation results displayed on the display 142. The evaluation results consist of a score (points) for each evaluation item and a short comment (message). (Table 1) TIFF2026005914000002.tif67138
[0034] For example, pronunciation errors can be displayed on the display 142 not only after the learning is completed but also during the learning process. In this case, the pronunciation errors can be highlighted in color or with an icon to draw the user's attention. Next, if necessary (such as when the user desires), the correct pronunciation for the user's incorrect pronunciation is output (step S200). Specifically, the central processing unit 121 starts a speech generation program (not shown), generates a correct pronunciation of the word that the user has mispronounced using synthesized speech, and outputs it from the speaker 143 . Furthermore, in addition to the evaluation of the user's language skills as described above, the foreign language acquisition assistance system 500 according to this embodiment also evaluates the user's non-verbal communication skills during conversation as follows.
[0035] When the user starts learning, camera 145 captures an image of the user (specifically, the user's non-verbal communication skills) during learning. The captured image data is recorded in user image data storage unit 132F via second program 131B. When the learning is completed (step S170), the central processing unit 121 starts the fourth program 131D, which compares the exemplary nonverbal communication skills during the conversation stored in the first data storage unit 132D with the nonverbal communication skills of the user during the conversation in the image data (image data capturing the nonverbal communication skills of the user during the conversation) captured by the camera 145 and sent to the control unit 120, via the learned second evaluation model stored in the second data storage unit 132E (step S210). As a result of this comparison, the central processing unit 121 generates an evaluation result that evaluates the user's non-verbal communication skills during conversation.
[0036] Table 2 shows an example of the evaluation (score and message) of facial expressions and gestures, which are non-verbal communication skills. (Table 2) TIFF2026005914000003.tif17119 After creating the evaluation results (Table 2), the central processing unit 121 displays the evaluation results together with the language skill evaluation results (Table 1) on the display 142 (step S220). The evaluation of the above non-verbal communication skills is detailed below.
[0037] The evaluation of the nonverbal communication skills (1) to (6) described above is performed by extracting the features of each of the user's nonverbal communication skills from the data of the user's image captured by the camera 145, and comparing these features with exemplary nonverbal communication skills during conversation stored in the first data storage unit 132D via a trained second evaluation model stored in the second data storage unit 132E. The feature amount is a numerical representation of the characteristics of each nonverbal communication skill in order to quantitatively evaluate the nonverbal communication skill. The fourth program 131D (evaluation means) compares the feature amount with exemplary nonverbal communication skills stored in the first data storage unit 132D. The feature amounts of each non-verbal communication skill are also used as training data 200 (see FIG. 3) for trained evaluation model 210B (this point will be described later).
[0038] (1)Facial expression Facial expressions are a major non-verbal communication skill that indicates emotions, intentions, attitudes, etc. In the foreign language learning assistance system 500 according to this embodiment, feature amounts for capturing various facial expression elements, centered on a smile, which is a particularly important element among facial expressions, are extracted from image data captured by the camera 145. (1A) Smile (1A-1) The degree of elevation of the corners of the mouth is determined by extracting the angle of the corners of the mouth, the distance of elevation, and symmetry as features. (1A-2) Regarding the mouth opening, the degree of mouth opening, the mouth opening area, and the mouth opening speed are extracted as feature quantities. (1A-3) Regarding the appearance of teeth, the degree of exposure of the upper teeth only, the upper and lower teeth, and the gums is extracted as a feature. (1A-4) Regarding cheek bulge, the contraction strength of the buccinator muscle, cheek bulge height, and bilateral symmetry are extracted as features. (1A-5) For wrinkles at the corners of the eyes, the number of wrinkles, depth, length, and symmetry are extracted as features.
[0039] (1B) Eyebrow movement Since eyebrow movements play an important role in expressing emotions and attitudes, the following feature amounts are extracted from the image data captured by the camera 145. (1B-1) Regarding the up and down movement of the eyebrows, the height of the eyebrows, the range of up and down movement, and the left-right symmetry are extracted as features. (1B-2) For wrinkles between the eyebrows, the contraction strength of the corrugator supercilii muscle, the distance between the eyebrows, and the depth of the wrinkles are extracted as features. (1B-3) The inclination angle of the arch of the eyebrows and symmetry are extracted as features of the inclination of the eyebrows. (1C) Eye opening The degree to which the eyes are open indicates the degree of surprise or interest, and the following feature amounts are extracted from the image data captured by the camera 145. (1C-1) The vertical distance of the palpebral fissure and the eye opening rate are extracted as features to represent the distance between the upper and lower eyelids. (1C-2) The exposed area of the iris and the exposed area of the sclera are extracted as features to indicate the degree of eye opening.
[0040] (1D) Blink Blinking indicates tension or the degree of concentration, and the following feature amounts are extracted from the image data captured by the camera 145. (1D-1) As the blink frequency, the number of blinks per unit time is extracted as a feature. (1D-2) The average duration of a single blink is extracted as a feature to determine the duration of the blink. (1D-3) The timing of left and right blinks is determined by extracting the time difference and asynchrony between the left and right blinks as features. (1E) Lip movement Lip movement is closely related to speech and emotional expression, and the following feature amounts are extracted from image data captured by the camera 145. (1E-1) Extract the mouth opening width, opening area, and opening speed as feature quantities to indicate the degree of mouth opening. (1E-2) The lip protrusion distance and left-right symmetry between the upper and lower lips are extracted as feature quantities to indicate the degree of lip protrusion.
[0041] (1E-3) To determine the degree of pulling in of the corners of the mouth, the amount of displacement of the corners of the mouth in the left-right direction and the curvature of the corners of the mouth are extracted as feature quantities. (1E-4) To measure lip tension, the degree of contraction of the muscles around the lips, such as the orbicularis oris, levator anguli oris, and depressor anguli oris, is extracted as a feature. (1F) Nose movement Nose movement expresses emotions such as disgust and anger, and the following feature amounts are extracted from image data captured by the camera 145. (1F-1) To determine the extent of the nostrils, the width of the nostrils, the area of the nostrils, and bilateral symmetry are extracted as features. (1F-2) To measure nasal wrinkles, the contraction strength of the proximal nasal muscle and the depth of the wrinkles on the upper part of the nose are extracted as features.
[0042] (1G) Cheek movement Cheek movements represent changes in emotions and facial expressions, and the following feature amounts are extracted from image data captured by the camera 145. (1G-1) Regarding cheek bulge, the contraction strength of the buccinator muscle, cheek protrusion height, and bilateral symmetry are extracted as features. (1G-2) To measure cheek tension, the degree of buccinator muscle tension and skin hardness are extracted as features. (1H) Jaw movement The movement of the jaw indicates confidence or tension, and the following feature amounts are extracted from the image data captured by the camera 145. (1H-1) As jaw protrusion, the horizontal movement distance of the tip of the chin and the angle of the temporomandibular joint are extracted as features. (1H-2) For chin pull, the tension of the muscles under the chin (such as the mentalis muscle) and the hardness of the skin are extracted as features.
[0043] (2) Gaze Gaze is important non-verbal information that indicates interest, attention, thought process, etc. in communication, and the following feature amounts related to gaze are extracted from image data captured by camera 145. (2A) Eye Contact Eye contact is used to evaluate the degree of interest and attention shown to the other person, and the following feature amounts are extracted from image data captured by the camera 145. (2A-1) As the frequency of eye contact, the number of times eye contact occurs per unit time is extracted as a feature from the image data captured by the camera 145. (2A-2) As the eye contact duration, the average duration of one eye contact is extracted as a feature from the image data captured by the camera 145. (2A-3) Types of eye contact, such as gazing at the other person's entire face, gazing at specific parts (eyes, mouth, nose, etc.), and avoiding gaze, are extracted as feature amounts from image data captured by camera 145.
[0044] (2B) Direction of line of sight The direction of the gaze is extracted from the image data captured by the camera 145. (2B-1) The horizontal direction of the gaze is determined by measuring the angle of the left and right eyeballs, and the deviation from the center of the face is calculated as a feature from the image data. (2B-2) The vertical direction of the gaze is measured by measuring the angle of the upper and lower eyeballs, and the deviation from the center of the face is calculated as a feature from the image data. (2C) Gaze duration The stability and concentration of the line of sight are extracted from image data captured by the camera 145. (2C-1) As the gaze time to a specific point, the gaze time to a specific point is measured as a feature from the image data. (2C-2) The gaze movement speed is calculated from the image data as a feature value, that is, the gaze movement angle per unit time, and the gaze movement is quantified.
[0045] (2D) Pupil diameter The size of the pupil is related to the degree of emotion or interest, and the size of the pupil is extracted from the image data captured by the camera 145. (2D-1) Pupil diameter: The diameter of the pupil is extracted as a feature from the image data. (2D-2) As for pupil contraction and dilation, the amount and speed of change in pupil diameter due to changes in lighting conditions and emotional changes are extracted as features from image data. (2E) Eye movement Eye movements are associated with thought processes and emotions, and are extracted from image data captured by the camera 145. (2E-1) Types of eye movements are classified into saccades (fast jumping movements), smooth pursuit movements (smooth tracking movements), and convergence (movement in which both eyes turn inward at the same time).
[0046] (2E-2) As for the eye movement speed, the speed of each eye movement is extracted from the image data as a feature. (2E-3) To determine the direction of eye movement, the horizontal and vertical components of each eye movement are extracted as features from the image data. (2E-4) As the eye movement frequency, the number of eye movements per unit time is extracted as a feature from the image data. (3) Gestures Gestures are body movements that convey information without using words and are an important non-verbal communication skill in language learning. The following gesture features are extracted from image data: (3A) Hand Movements Hand movements play an important role in expressing emotions and transmitting information, and hand movements are extracted as feature amounts from image data captured by the camera 145.
[0047] (3A-1) The hand position is measured in three dimensions, relative to the body parts (e.g., head, shoulders, torso) for each frame, and the trajectory of the hand movement is extracted as time-series data (features). (3A-2) As the speed of hand movement, the distance traveled by the hand per unit time and the acceleration are measured, and the speed of hand movement is extracted and quantified as a feature. (3A-3) The magnitude of hand movement is measured by analyzing the range of hand movement and the complexity of the trajectory, and the dynamism of the movement is extracted as a feature. (3B) Pointing Pointing is a type of gesture used to indicate a specific object or direction, and the pointing is extracted from image data captured by the camera 145. (3B-1) The direction of the pointing finger is determined by extracting the directional vector and horizontal and vertical angles of the fingertip as features from the image data (this makes it possible to identify the target of the pointing finger). (3B-2) As the frequency of pointing, the number of times pointing occurs per unit time is extracted as a feature from the image data. (3B-3) As the duration of the pointing, the average duration of one pointing is extracted as a feature from the image data.
[0048] (3C) Palm orientation The orientation of the palm represents emotions and attitudes, and the orientation of the palm is extracted as a feature from image data captured by the camera 145. (3C-1) To determine whether the palm is facing up or down, the angle between the normal vector of the palm and the vertical axis is extracted as a feature from the image data. (3C-2) To determine whether the palm is facing forward or backward, the angle between the normal vector of the palm and the line of sight is extracted as a feature from the image data. (3C-3) To determine whether the palm is facing inward or outward from the body, the angle between the normal vector of the palm and the left-right axis is extracted as a feature from the image data.
[0049] (3D)Applause Applause expresses joy or approval, and the following items are extracted from the image data captured by the camera 145. (3D-1) As the speed of clapping, the number of clappings per unit time is extracted as a feature from the image data captured by the camera 145. (3D-2) To determine the strength of clapping, the acceleration of the hands and the sound pressure level during clapping are extracted as feature quantities from the image data captured by the camera 145 or the sound data acquired by the microphone 144 . (3D-3) As the duration of clapping, the average duration of one clap is extracted as a feature from the image data captured by the camera 145. (3E) Hand shape Hand shapes express emotions and attitudes, and the following items are extracted as feature quantities from image data captured by the camera 145. (3E-1) As the degree of hand opening and closing, the degree of finger opening and the distance between the fingers are extracted as feature values from the image data captured by the camera 145, and the degree of hand opening is quantified.
[0050] (3E-2) Regarding the hand grip, the shape of the hand, such as a clenched fist, a light clenched fist, or fingers extended, is extracted as a feature from the image data captured by the camera 145. (3E-3) Regarding hand clasping, the way the hands are clasped, such as clasping fingers or placing hands on top of each other, is extracted as a feature from the image data captured by the camera 145. (3F) Arm Movement Arm movements complement explanations and emotional expressions, and the following items are extracted as feature amounts from image data captured by the camera 145. (3F-1) The arm position is measured in three dimensions, relative to the body parts (e.g., head, shoulders, torso) for each frame, and the trajectory of the arm movement is extracted as time-series data (features). (3F-2) As the arm movement speed, the average arm movement distance and average acceleration per unit time are extracted as feature quantities from the image data captured by the camera 145, and the speed of the arm movement is quantified. (3F-3) The magnitude of arm movement, such as the range of arm movement and the complexity of the trajectory, is extracted as features from the image data to evaluate the dynamism of the movement.
[0051] (3G) Elbow movement forms a gesture in conjunction with arm movement, and the following items are extracted as feature amounts from image data captured by camera 145. (3G-1) Regarding elbow bending and straightening, the angle of the elbow joint is extracted as a feature from the image data captured by the camera 145, and the degree of elbow bending and straightening is quantified. (3G-2) As for the rotation of the elbow, the internal rotation and external rotation angles of the elbow joint are extracted as feature values from the image data captured by the camera 145, and the rotational movement of the elbow is quantified. (3H) Shoulder movements express emotions and attitudes, and the following items are extracted as feature amounts from image data captured by the camera 145. (3H-1) As for the up and down of the shoulders, the up and down movement distance of the shoulder blades is extracted as a feature amount from the image data captured by the camera 145. (3H-2) Regarding the anterior-posterior movement of the shoulder, the anterior-posterior movement distance of the scapula is extracted as a feature amount from the image data captured by the camera 145. (3H-3) Regarding shoulder rotation, the internal and external rotation angles of the shoulder joint are extracted as feature quantities from the image data captured by the camera 145.
[0052] (4) Body language Body language is non-verbal information that conveys emotions, attitudes, confidence, etc. through posture, actions, bodily movements, etc. The following body language features are extracted from image data: (4A) Posture (standing) (4A-1) As the foot width, the distance between both feet and the positional relationship between the left and right feet (parallel, V-shaped, reverse V-shaped, etc.) are extracted as feature amounts from the image data captured by the camera 145. (4A-2) As for the center of gravity position, the deviation of the center of gravity position from the left and right and the front and rear is extracted as a feature amount from the image data captured by the camera 145. (4A-3) The degree of S-curve of the spine, hunched back, arched back, etc. are extracted as feature quantities from the image data captured by the camera 145 as curvature of the spine. (4A-4) Regarding the position of the scapula, the degree of adduction / abduction and upward / downward rotation of the scapula is extracted as feature amounts from the image data captured by the camera 145. (4A-5) The tilt angle of the head, forward, backward, left and right, is extracted as a feature from the image data captured by the camera 145 as the tilt of the head.
[0053] (4B) Posture (sitting posture) (4B-1) To determine the degree of leaning against the backrest, the contact area with the backrest and the inclination angle of the body are extracted as feature quantities from the image data captured by the camera 145. (4B-2) Regarding the manner in which the legs are crossed, whether the legs are crossed or not and which leg is on top are extracted as feature amounts from the image data captured by the camera 145. (4B-3) Regarding the position of the arms, whether the arms are folded, open, placed at the side of the body, or placed on a desk is extracted as a feature from the image data captured by the camera 145. (4B-4) Regarding the way the hands are placed, whether the hands are placed on the lap, clasped, or clenched, is extracted as a feature from the image data captured by the camera 145.
[0054] (4C) Head movement (nodding) (4C-1) As the nodding speed, the average number of noddings per unit time is extracted as a feature from the image data captured by the camera 145. (4C-2) As the nodding angle, the maximum and minimum angles of the nodding are extracted as feature amounts from the image data captured by the camera 145. (4C-3) As the nodding frequency, the number of nods is extracted as a feature from the image data captured by the camera 145. (4D) Head movement (head shaking) (4D-1) As the swing speed, the average number of swings per unit time is extracted as a feature from the image data captured by the camera 145. (4D-2) As the swing angle, the maximum swing angle and the minimum swing angle are extracted as feature amounts from the image data captured by the camera 145. (4D-3) As the head shaking frequency, the number of head shakings is extracted as a feature from the image data captured by the camera 145.
[0055] (4E) Body Orientation (4E-1) As the degree of facing the other person, the angle between the front of the body and the other person is extracted as a feature from the image data captured by the camera 145. (4E-2) The degree of opening and closing of the body, such as the degree to which the arms and legs are open and the proportion of time the body is facing forward, are extracted as feature amounts from the image data captured by the camera 145. (4F) Walking Directions (4F-1) As the stride length, the average walking distance per step is extracted as a feature from the image data captured by the camera 145. (4F-2) As the walking speed, the average walking distance per unit time is extracted as a feature from the image data captured by the camera 145. (4F-3) Regarding arm swing, the magnitude of arm swing, left-right symmetry, and the degree of agreement between arm swing and gait are extracted as feature amounts from the image data captured by the camera 145.
[0056] (4G) Body tilt (4G-1) As the degree of forward or backward leaning, the relative position of the center of gravity of the body and the center of the sole of the foot, and the inclination angle of the trunk are extracted as feature quantities from the image data captured by the camera 145. (4G-2) As for the inclination to the left and right, the shift of the center of gravity of the body to the left and right, and the angle of inclination of the trunk to the left and right are extracted as feature amounts from the image data captured by the camera 145. (5) Proxemics Proxemics is non-verbal information that indicates spatial relationships such as distance, angle, and space occupation in interpersonal space. The following proxemics-related features are extracted from image data. (5A) Personal Space The sense of distance relative to a fixed camera position is extracted as a feature amount from image data captured by the camera 145. Specifically, the change in sense of distance within a frame is analyzed as a change in distance.
[0057] (5B) Interpersonal angle The angle of the body relative to the camera 145 is extracted as a feature amount from the image data captured by the camera 145. Specifically, the angle of the body is observed as the frontality relative to the camera and changes in the angle. (5C)Territoriality The awareness of spatial occupation and the strength of self-assertion are evaluated from the movements and postures. The following items are extracted as feature quantities from the image data captured by the camera 145. (5C-1) As for space occupation, the method of occupying a fixed space is extracted as a feature from the image data captured by the camera 145. (5C-2) Regarding physical barriers, the presence or absence and degree of physical barriers such as crossed arms or crossed legs are extracted as feature quantities from the image data captured by the camera 145.
[0058] (6) Physical Appearance Physical appearance (including grooming) is non-verbal information that significantly influences first impressions and self-expression. In particular, physical appearance has a significant impact on the effectiveness of communication in business situations such as interviews and presentations. The following physical appearance-related features are extracted from image data and scored based on evaluation criteria. (6A)Clothing (6A-1) Color and Design: Color selection and design are evaluated for personal impression and appropriateness for TPO (Time, Place, Occasion), and are extracted as feature amounts from image data captured by the camera 145. (6A-2) Regarding cleanliness and appropriateness, the cleanliness of the clothes and their appropriateness according to the time, place, and occasion are extracted as feature quantities from the image data captured by the camera 145.
[0059] (6B) Hairstyle (6B-1) Regarding hairstyle and hair color, the influence that the choice of hairstyle and hair color has on the personal impression and the time, place, and occasion is extracted as a feature from the image data captured by the camera 145. (6B-2) Regarding cleanliness, the cleanliness and grooming status of the hair are extracted as features from the image data captured by the camera 145. (6C) Accessories The type of accessory and its appropriateness for the occasion are extracted as feature amounts from image data captured by the camera 145. (6D) Appearance (6D-1) Regarding the cleanliness and care of the skin, the state of care of the skin, nails, and beard is extracted as a feature from the image data captured by the camera 145. (6D-2) Regarding grooming, the cleanliness and the level of care are extracted as features from the image data captured by the camera 145.
[0060] (6E) Posture The appropriateness of posture is evaluated by evaluating the correctness of the standing and sitting posture. Specifically, whether the back is straight, whether a natural posture is maintained, etc. are extracted as feature amounts from image data captured by the camera 145. As described above, the feature quantities extracted in this manner can be used as input 220 for constructing and updating the trained evaluation model 210B. An example of the construction and training (updating) of the evaluation model 210B will be described below. For example, the evaluation model 210B is constructed as a multimodal deep learning model and a large-scale language model (LLM) using supervised machine learning. Supervised machine learning is a method of training an evaluation model using pairs of input data and correct labels. In this example, video data in a wide variety of languages is used as input, and the evaluation scores for each evaluation item of non-verbal communication skills are treated as correct labels.
[0061] For example, the following features are used as input data: (a) Facial expression features For example, features of facial expressions include the strength of Action Units (AUs) based on the Facial Action Coding System (FACS), coordinates of facial landmarks (eyes, nose, mouth, etc.), and facial shape features. (b) Head posture For example, the three-dimensional rotation angle of the head (pitch angle, yaw angle, roll angle, etc.), the head movement speed, the head movement trajectory, etc. are used as head posture feature quantities. (c) Audio features For example, Mel-frequency cepstral coefficients (MFCC), linear predictive coding (LPC) coefficients, formant frequencies, fundamental frequencies (F0), etc. are used as speech features.
[0062] (d) Prosodic features For example, the features of prosodic features include the pitch, volume, speech rate, rhythm, and time-series data of intonation. The output data is, for example, a score (for example, on a five-point scale from 1 point to 5 points) for each evaluation item of non-verbal communication skills. The number of nodes in the output layer corresponds to the number of types of nonverbal communication skills, and each node represents an evaluation score for the evaluation item of the corresponding nonverbal communication skill. Deep learning architectures such as convolutional neural networks (CNNs), LSTMs (Long Short-Term Memory), and Transformers are used as learning algorithms.
[0063] Convolutional neural networks are suitable for extracting features from images and time-series data, and are particularly effective for facial expression and gesture recognition. LSTM is a type of recurrent neural network suitable for handling time series data, and is particularly effective for recognizing time series patterns of speech and prosody. Transformer uses a self-attention mechanism and is suitable for parallel processing, especially when handling multiple non-verbal information in an integrated manner. The loss functions used are mean squared error (MSE) for regression problems and cross entropy error for classification problems. The mean squared error (MSE) is used to predict the continuous score for each assessment item of nonverbal communication skills, and the cross-entropy error is used to determine whether a particular nonverbal communication skill belongs to a particular category (e.g., whether the user's facial expression is smiling or not).
[0064] The optimization algorithm used is Adam (Adaptive Moment Estimation). Generally, Adam is used as the initial setting because it has fast learning convergence and it is relatively easy to adjust hyperparameters. If Adam's performance is not sufficient, SGD (Stochastic Gradient Descent) or RMSprop (Root Mean Square Propagation) can also be used. As mentioned above, nonverbal communication skills span multiple modalities (types of information), such as facial expressions, gaze, gestures, body language, and hand movements. In order to comprehensively analyze and evaluate these diverse nonverbal communication skills, it is effective to construct the evaluation model 210B as a multimodal deep learning model.
[0065] Below, we provide an overview of multimodal deep learning models. In the multimodal deep learning model, the following features are input 220 (see FIG. 3): (1) Facial information Using algorithms such as Haar Cascades, HOG+SVM, and MTCNN, the user's face area is detected from the user's video recorded in the user image data storage unit 132F. Next, using libraries such as Dlib, OpenFace, and MediaPipe Face Mesh, landmarks such as 68, 128, and 468 points on the user's face are detected. The detected landmarks are used as features that represent changes in facial shape and facial expressions. Using facial landmark information, texture information, or pre-trained models (e.g., VGGFace, FaceNet), basic emotions (happiness, sadness, anger, fear, surprise, disgust) and facial expressions (confusion, contempt, interest, etc.) are recognized. To capture subtle changes in emotions and facial expressions, the movement of each facial muscle is expressed as an Action Unit (AU) based on the FACS (Facial Action Coding System), and its strength is analyzed.
[0066] (2) Gaze information The gaze direction is estimated based on facial landmark information and pupil position information. Furthermore, gaze movements (saccades, smooth pursuit movements, etc.), gaze duration, and changes in pupil diameter are recorded as time-series data and extracted as feature quantities. (3) Audio information Using a VAD (Voice Activity Detection) algorithm, silent periods are removed from the user's voice signal recorded in the user voice storage unit 132B, and only speech periods are extracted. Next, acoustic features such as MFCC (Mel-Frequency Cepstral Coefficients), LPC (Linear Predictive Coding), formant frequencies, and fundamental frequency (F0) are extracted. Furthermore, we use tools such as prosodylab-aligner to extract prosodic features such as pitch contour, intensity, and duration.
[0067] (4) Physical information Using pose estimation algorithms such as OpenPose, AlphaPose, and PoseNet, the user's skeletal information is extracted and time series data on the positions and angles of the body's joints is obtained. Next, feature quantities such as the magnitude, speed, direction, and joint angles of the user's body movements are calculated from the skeletal information. (5) Linguistic information A newly machine-learned large-scale language model (LLM) is used to perform grammatical, lexical, and semantic analysis of speech content. It analyzes the emotions and intentions of users based on what they say, as well as the context and intention of their language, and uses this to interpret their non-verbal communication skills. As described above, after the feature amounts are extracted from each modality, preprocessing such as standardization and normalization is performed to unify the scale of the feature amounts.
[0068] Specifically, the learning efficiency and generalization performance of the model are improved by removing redundant features and features that become noise. Furthermore, modalities are fused. For example, an attention mechanism is used to learn the associations between features of different modalities and dynamically adjust the importance of each modality depending on the situation. Then, tensor fusion is used to represent the features of different modalities as tensors, and the modalities are integrated using techniques such as tensor decomposition. In this case, self-attention and location encoding can be used to efficiently process multimodal data, including time series and spatial data. Based on the integrated features, an evaluation score for each non-verbal communication skill is calculated (output 230). For the evaluation, a regression model, a classification model, or reinforcement learning can be used.
[0069] For example, regression models can be used that predict continuous values, such as linear regression, support vector regression (SVR), random forest regression, etc. Classification models can be used that perform category classification, such as logistic regression, support vector machine (SVM), decision tree, random forest, etc. According to the foreign language learning assistance system 500 of this embodiment, the following effects (merits) can be obtained. The foreign language learning assistance system 500 enables the user to improve not only the language skills of the user but also the non-verbal communication skills of the user. Specifically, the system enables the user to self-study exemplary non-verbal communication skills during conversation, thereby improving the user's communication ability. As shown in this embodiment, the foreign language learning assistance system 500 can be built into a currently widely used mobile wireless communication device 100 such as a smartphone. Typically, a mobile wireless communication device 100 such as a smartphone is always carried nearby, allowing the user to study a foreign language (especially non-verbal communication skills) at their own convenience, regardless of time or place.
[0070] (Second embodiment) There is a close relationship between what an interlocutor says and their nonverbal communication skills, and the evaluation of their nonverbal communication skills may change depending on what the interlocutor says. For example, generally speaking, a smile and a bright tone of voice are appropriate when expressing joy or surprise, while a calm tone of voice and modest gestures are required when expressing sadness. Conversely, a sad expression or a gloomy voice in a happy situation, or a smile or a calm voice in an angry situation, will appear extremely unnatural to the other person. Therefore, the fourth program 131D (evaluation means) in the foreign language learning assistance system according to this embodiment uses the content of what the user has said as language data in an auxiliary manner when evaluating the user's non-verbal communication skills.
[0071] The following items are used as language data: (a) Speech content: Evaluate the grammar, vocabulary, and pronunciation of what is said. (b) Context: Analyze the context and intention of the utterance to help interpret nonverbal communication skills. (c) Phrase use: Evaluate the frequency and appropriateness of specific phrases and expressions and consider their relationship to non-verbal communication skills. (d) Speech Fluency: Evaluates how smoothly the speech is delivered and analyzes the consistency of non-verbal communication skills. The fourth program 131D (evaluation means) analyzes each item of these language data (for example, extracts the above-mentioned feature amounts) and reflects the analysis results in the evaluation of the user's non-verbal communication skills.
[0072] For example, in the example mentioned above, if the user shows an irritated expression or turns away when the topic is about something that should be happy, the fourth program 131D (evaluation means) will give the user a low rating for their non-verbal communication skills (for example, in terms of facial expressions and gestures) when the user behaves in this way, even though the content of the conversation indicates that the situation should be happy. In this way, by taking into account not only the image data captured by camera 145 but also the user's language data recorded in user voice storage unit 132C, it becomes possible to more accurately evaluate the user's non-verbal communication skills. Additionally, the following items can be used as auxiliary language data:
[0073] (a) Paralanguage Paralanguage is a non-verbal element contained in a speaker's voice and speaking style, and plays an important role in conveying emotions, attitudes, intentions, etc. The following paralanguage-related features are extracted from speech data and used as auxiliary data. (A) Tone of voice Voice tone is an important factor in expressing emotions and attitudes, and the following features related to voice tone are extracted from the voice data. (A-1) The average value, standard deviation, and range of variation of the fundamental frequency are extracted as features from the voice data to indicate the pitch of the voice, and finally the pattern of change in pitch of the voice is extracted. (A-2) To determine the brightness of the voice, features such as the height of the spectral center of gravity and the proportion of high-frequency components are extracted from the voice data, and ultimately the brightness and softness of the voice are extracted.
[0074] (A-3) To measure the strength of the voice, sound pressure level and voice energy are extracted as features from the voice data, and ultimately the strength and power of the voice are extracted. (A-4) To measure voice stability, the stability of fundamental frequency and volume is extracted as features from the voice data, and finally, voice tremor and instability are extracted. (B) Pitch of voice Voice pitch is an important factor in conveying emotions and the intention of speech, and the following features related to voice pitch are extracted from speech data. (B-1) Fundamental frequency features such as the average fundamental frequency, highest fundamental frequency, and lowest fundamental frequency during speech are extracted from the voice data, and ultimately the overall characteristics of the voice pitch are extracted. (B-2) Pitch changes such as the range of pitch fluctuations, pitch fluctuation patterns (rising and falling), and frequency of rising and falling are extracted as features from the voice data, and ultimately vocal intonation and emotional expression are extracted.
[0075] (C) Volume of voice The volume of a voice is an important factor that indicates ease of listening and the degree of confidence, and the following features related to voice volume are extracted from the voice data. (i) C-1) Regarding volume, the average, maximum, and minimum sound pressure levels are extracted as features from the audio data, and ultimately the overall characteristics of the voice loudness are extracted. (iC-2) The distance the voice can reach is estimated as a feature from the voice data by comparing the voice attenuation characteristics and the surrounding noise level. (I) Speaking speed Speaking speed is an important factor that reflects the ease of listening and the speaker's personality, and the following features related to speaking speed are extracted from the speech data. (iD-1) To measure speaking speed, the average number of words spoken per minute and the average number of syllables are counted as features from the audio data.
[0076] (iD-2) As for the intervals between speeches, the length of the silent periods (pauses) between utterances is measured as a feature from the audio data, and the intervals between speeches are analyzed. (E) Intonation Intonation plays an important role in conveying emotions and speech nuances, and the following features related to intonation are extracted from speech data. (E-1) Intonation, such as rising and falling pitch and intonation at the end of a sentence, is extracted as features from the speech data, and ultimately the rhythm and emotion of the entire speech are extracted. (E-2) Accent, such as the position and degree of stress in words and phrases, is extracted as a feature from the audio data, and ultimately the clarity and ease of listening of the speech is extracted. (E-3) Rhythm: The rhythmic pattern, regularity, tempo, etc. of speech are extracted as features from the audio data, and ultimately the fluency and naturalness of the speaking style are extracted.
[0077] (I F) Sound volume The strength of a sound is an important element used for emphasis and emotional expression, and the following features related to strength of a sound are extracted from the speech data. (iF-1) When adding emphasis, the words or phrases to be emphasized, the degree of change in emphasis, etc. are extracted as features from the audio data, and it is then determined whether the emphasis has been effective. (I F-2) To provide contrast, differences in volume and richness of intonation are extracted as features from the audio data, and ultimately expressiveness and ease of listening are extracted. (IG) Voice quality Voice quality is an important factor that influences ease of listening and impression, and the following features related to sound quality are extracted from the voice data. (I G-1) The clarity of the voice is determined by extracting features from the voice data, such as the proportion of high frequency components and the lack of noise, and ultimately extracting the clarity and ease of listening of the voice. (I G-2) The richness of vocal cord vibration and the degree of resonance are extracted as features from the audio data to represent the sound of the voice, and ultimately the depth and richness of the voice are extracted.
[0078] (I G-3) The roughness of the voice is determined by extracting features from the voice data, such as the proportion of low-frequency components and the amount of noise, and ultimately extracting the roughness and difficulty of listening of the voice. (IH) Speech clarity The clarity of speech is an important factor that influences the smoothness of communication, and the following features related to speech clarity are extracted from speech data. (i) H-1) To assess the accuracy of pronunciation, features such as the accuracy of vowels, consonants, and intonation are extracted from the speech data, and the clarity of the pronunciation is ultimately determined. (i) H-2) Regarding articulation, the clarity and fluency of pronunciation are extracted as features from the speech data. (I) Filler Fillers are words such as "ah" and "eh" that are uttered unconsciously during speech, and excessive use of them can make speech difficult to understand. The following filler-related features are extracted from speech data.
[0079] (I-1) The average number of times fillers such as "ah" and "eh" are used per unit time is extracted as a feature from the speech data. (I-2) Types of fillers such as "ah," "eh," "um," and "uh," are classified as features from the speech data. (b) Silence Silence refers to a state in which there is no speech. Silence is non-verbal information that can have various meanings, such as a pause in conversation, time for thinking, or an expression of emotion. The following features related to silence are extracted from speech data. (B) Intentional silence (a) A-1) The length of the silent intervals (pauses) between utterances and the context before and after the pauses are extracted as features from the audio data to determine the appropriateness of the spacing. (a)-2) The length of silence is extracted as a feature from the audio data, such as less than 1 second, 1-2 seconds, 2-3 seconds, or more than 3 seconds, and it is finally determined whether the length of silence is appropriate for the situation. (a)-3) As the frequency of silence, the average number of silences per unit time is extracted as a feature from the speech data, and the frequency of silence is finally determined.
[0080] (B) Between conversations (B-1) As for the length of a pause, the length of a short pause during speech is extracted as a feature from the speech data. (B-2) As the frequency of pauses, the number of pauses per unit time is extracted as a feature from the speech data. (B-3) Regarding how to take a pause, the context before and after the pause is extracted as features from the audio data, and the pause is ultimately evaluated as being natural, unnatural, or drawn-out. (B) Response (B-C-1) As for the timing of backchanneling, the timing of backchanneling in response to the other person's speech is extracted as a feature from the voice data. (C-2) Types of interjections such as "Yes," "Yeah," "Uh-huh," and "I see" are classified from the voice data. (C-3) As the frequency of backchannel responses, the average number of backchannel responses per unit time is extracted as a feature from the speech data.
[0081] (c) How to use time Chronemics is non-verbal information that indicates attitudes and behaviors toward time, and varies significantly depending on culture and situation. The following features related to time usage are extracted from speech data. (HaA) Reaction time (HaA-1) To measure the speed of response to a question, the time from when the question is asked to when the answer begins is extracted as a feature from the voice data. (HaA-2) As the speed of reaction to an action, the time from receiving an instruction to take an action to starting the action is extracted as a feature from the voice data or image data.
[0082] (HaB) Pace of conversation (B-1) The speed of speaker change is measured by extracting a feature from the speech data, which is the time it takes for one speaker to take over from another speaker. (HaB-2) The speed at which a topic develops is determined by extracting features from the speech data, such as the time it takes to talk about one topic and how often the topic changes. (HaC) Interval of conversation Silent periods between utterances, such as pauses in the middle of a sentence or silent periods between words, are extracted as features from the audio data and evaluated to determine whether they are natural intervals. (H D) Auditory information (HaD-1) Voice quality, such as voice vitality, breathing, and the presence or absence of nasality, are extracted as features from the voice data. (HaD-2) Regarding laughter, the frequency, volume, and type of laughter (for example, loud laughter, chuckling, etc.) are extracted as features from the audio data. (HaD-3) As for throat clearing, the frequency and timing of throat clearing are extracted as features from the voice data.
[0083] (Third embodiment) In the foreign language acquisition support system 500 according to the first embodiment, evaluation results are created for the user's language skills and non-verbal communication skills, but this can be further developed to create a curriculum specialized for the user in order to compensate for the language skills and non-verbal communication skills of a user who received low evaluations in the evaluation results. The foreign language acquisition support system according to the third embodiment has a sixth program (not shown) in the application section 131 of the external memory 130, and this sixth program functions as a curriculum creation means for creating a curriculum (learning plan) indicating future learning guidelines in accordance with the user's evaluation results. When the user evaluation results (steps S180 and S210 in FIG. 4) are generated, the central processing unit 121 starts a sixth program.
[0084] A database (not shown) listing countermeasures for each shortcoming in language skills and non-verbal communication skills is stored in advance in the data storage unit 132, and the sixth program searches this database for countermeasures to address the shortcomings pointed out in the evaluation results, and compiles these countermeasures into a user-specific curriculum for each of language skills and non-verbal communication skills. Table 3 is an example of a curriculum (learning plan) for language skills. (Table 3) TIFF2026005914000004.tif62139
[0085] Table 4 is an example of a curriculum (learning plan) for nonverbal communication skills. (Table 4) TIFF2026005914000005.tif62160 These curricula are displayed on the display 142. The foreign language acquisition assistance system according to this embodiment allows users to obtain a curriculum (study plan) that is tailored to their weaknesses in both language skills and non-verbal communication skills. This curriculum also provides practice sessions that focus on the user's weaknesses, allowing the user to efficiently overcome their weaknesses.
[0086] (Fourth embodiment) The foreign language acquisition support system according to the fourth embodiment stores a seventh program (not shown) in the application section 131 of the external memory 130, and this seventh program has the function of creating new learning materials that allow the user to study according to the curriculum created in the third embodiment. In other words, the seventh program creates new learning materials that are tailored to the user's weaknesses identified in the evaluation results. The learning materials are configured as, for example, images (still images) consisting of learning text (subtitles) and audio. The central processing unit 121 starts the seventh program and creates an image using the image creation technique. The created image is stored in, for example, the sub-data storage unit 132G. Sentence generation can be performed using, for example, LLMs (Large Language Models) and RAG (Retrieval-Augmented Generation). In this way, the foreign language learning assistance system according to this embodiment allows users to obtain new learning materials that are specialized to address their own weaknesses, thereby improving learning efficiency.
[0087] (Fifth embodiment) In the fourth embodiment, images (still images) are created as learning materials, but it is also possible to create a video in which the learning sentences are reproduced as audio. Furthermore, in addition to the audio, the video can also display the sentences as subtitles. The foreign language acquisition support system according to the fifth embodiment stores an eighth program (not shown) in the application section 131 of the external memory 130, and this eighth program has the function of creating new videos as learning materials that allow users to study according to the curriculum created in the third embodiment. The central processing unit 121 starts the eighth program and creates a moving image according to the specified conditions using image creation technology and voice synthesis technology. The created moving image is stored in the moving image storage unit 132A. For example, this video can be created as an interactive video in which the user and a person (character) in the video have a conversation.
[0088] FIG. 5 is a conceptual diagram of a foreign language learning assistance system according to this embodiment. As shown in FIG. 5, a character 310 appears on the screen of the display 142, and a user 320 faces the screen of the display 142 and interacts with the character 310. The foreign language learning assistance system of this embodiment stores ninth, tenth, and eleventh programs (none of which are shown in the figure) in the application section 131 of the external memory 130, and the ninth program has the function of creating a response to speech uttered by the user 320 and converting it into speech, the tenth program has the function of selecting an appropriate non-verbal communication skill corresponding to the non-verbal communication skill of the user 320, and the eleventh program has the function of creating subtitles in accordance with instructions from the central processing unit 121 and displaying the subtitles on the screen of the display 142.
[0089] When the user 320 starts a conversation with the character 310, the user 320's voice is collected by the microphone 144 and sent to the central processing unit 121 as voice data. The central processing unit 121, which has received the voice data, starts the ninth program, creates a response to the voice uttered by the user 320, and vocalizes it. The vocalized response is output from the speaker 143 as a voice uttered by the character 310. In this way, a dialogue is established vocally between the character 310 on the display 142 and the user 320. Furthermore, when the user 320 starts a conversation, the appearance and actions (non-verbal communication skills) of the user 320 are captured by the camera 145, and the captured image is sent to the central processing unit 121 as image data. The central processing unit 121 that has received the image data starts the eighth program and uses image creation technology to create an animation in which the character 310 moves in a way that matches the movements of the user 320. The created animation is displayed on the display 142. In this way, a visual interaction is established between the character 310 on the display 142 and the user 320.
[0090] Upon receiving the image data, the central processing unit 121 executes the ninth and tenth programs (which create audio and visual dialogue between the character 310 and the user 320) and simultaneously starts the fourth program 131D. This causes the exemplary non-verbal communication skills during conversation stored in the first data storage unit 132D to be compared with the non-verbal communication skills during conversation of the user 320 in the image data of the user 320 via the trained second evaluation model 210B stored in the second data storage unit 132E. As a result of this comparison, the central processing unit 121 generates an evaluation result (see Table 2) for the user's 320 non-verbal communication skills during conversation. Next, the central processing unit 121 starts the eleventh program, creates subtitles that reflect the evaluation results, and displays the subtitles 330 (see FIG. 5) on the screen of the display 142. In this way, in the foreign language acquisition support system according to this embodiment, if a non-verbal communication skill of the user 320 that is rated low appears while the user 320 is conversing with the character 310, subtitles 330 of that content appear on the display 142.
[0091] Therefore, the user 320 can know in real time when his / her own shortcomings in non-verbal communication skills are appearing and can respond immediately, thereby improving learning efficiency. For example, if the user 320 does not move his / her face during a conversation, the subtitles 330 may indicate that "his / her face has no expression." Furthermore, following such a comment, advice can be given. For example, following the comment "Your face is expressionless," you can give advice such as "Smile more." This allows the user 320 to know his / her own shortcomings and understand in real time how to deal with them, thereby promoting the improvement of non-verbal communication skills. Also, instead of the subtitles 330 displayed on the screen of the display 142, or together with the subtitles 330, it is possible to configure the character 310 to speak the content of the subtitles 330.
[0092] For example, the foreign language learning assistance system further includes a twelfth program (not shown), which functions as a subtitle voice conversion means for converting the subtitles 330 into voice. The twelfth program plays a voice that verbalizes the subtitles 330 as the words of the character 310 on the screen of the display 142 instead of the subtitles 330 or together with the subtitles 330. Generally, humans can understand things more quickly through hearing than through sight. Therefore, compared to using only subtitles 330, the user 320 can easily and immediately recognize the shortcomings of his / her non-verbal communication skills by being directly pointed out verbally by the character 310, who is the other party in the dialogue. This is particularly effective when the user 320 does not have time to read the subtitles 330.
[0093] (Sixth embodiment) Unlike verbal communication, nonverbal communication can vary greatly depending on cultural background (including religious background), as the same nonverbal communication skills may have different meanings in different cultural spheres. For example, the appropriate frequency and duration of eye contact differs depending on the culture, and some gestures are only understood in certain cultures. The foreign language learning assistance system according to the sixth embodiment aims to evaluate non-verbal communication skills appropriately for users from diverse cultural backgrounds, taking into consideration the cultural backgrounds of the user and the other party. For this reason, the foreign language acquisition support system of this embodiment includes a cultural background information database (not shown) that stores information on the cultural background collected from each country or region, and a thirteenth program (not shown). The cultural background information database is stored in the data storage unit 132, and the thirteenth program is stored in the application storage unit 131.
[0094] The thirteenth program functions as an adjustment means that reads out the user's cultural background information from the cultural background information database and makes adjustments to the evaluation of the user's non-verbal communication skills by the fourth program 131D taking into account the user's cultural background information. The user's cultural background is either specified by the user before starting foreign language learning, or determined by using an algorithm that adapts to the user's cultural background in real time during the assessment. For example, in one culture A, constantly smiling is considered a healthy non-verbal communication skill, but in another culture B, smiling may be considered a mockery of the other person. For this reason, when a user belonging to culture A is dealing with a person belonging to culture B, the thirteenth program (adjustment means) adjusts the evaluation of the user's non-verbal communication skills by the fourth program 131D so as to give a lower rating to the user's smiling. In addition, in the interactive video shown in the fifth embodiment, it is also possible to display subtitles 330 advising the user 320 not to smile, or to have the character 310 point this out verbally. [Explanation of symbols]
[0095] 500 Foreign language acquisition support system according to the first embodiment of the present invention 100 Portable wireless communication device 120 control section 121 Central Processing Unit 130 external memory 131 Application Storage 132 Data storage unit 142 displays 145 Camera 210 Evaluation Model 310 characters 320 users 330 subtitles
Claims
1. A foreign language acquisition support system that supports a user in acquiring a foreign language, an imaging means for imaging the user making sounds in accordance with the audio of the video; a first storage means for storing exemplary non-verbal communication skills of interlocutors during a conversation; a second storage means for storing a trained evaluation model for evaluating non-verbal communication skills of interlocutors during a conversation; evaluation means for comparing the exemplary nonverbal communication skills stored in the first storage means with the nonverbal communication skills of the user captured by the imaging means using the trained evaluation model stored in the second storage means, and evaluating the nonverbal communication skills of the user; a display means for displaying the evaluation created by the evaluation means; A foreign language acquisition support system equipped with:
2. 2. The foreign language acquisition support system according to claim 1, wherein the evaluation means evaluates the user's non-verbal communication skills based on at least one evaluation item of facial expressions, gaze, gestures, body language, proxemics, physical appearance, visual focus, auditory information, and cultural background.
3. The foreign language acquisition support system according to claim 2, characterized in that the evaluation means extracts features quantitatively indicating the characteristics of each of the evaluation items from the data of the image of the user captured by the imaging means, and compares these features with exemplary non-verbal communication skills stored in the first storage means.
4. 2. The foreign language acquisition support system according to claim 1, wherein the evaluation means uses the content of the user's conversation as auxiliary language data when evaluating the user's non-verbal communication skills.
5. a cultural background information database that stores cultural background information for each country or region; an adjustment means for reading out cultural background information of the user from the cultural background information database and adjusting the evaluation of the user's non-verbal communication skills by the evaluation means in consideration of the cultural background information of the user; 2. The foreign language learning assistance system according to claim 1, further comprising:
6. the trained evaluation model is generated by machine learning to evaluate the non-verbal communication skills of the interlocutor using exemplary non-verbal communication skills stored in the first storage means as a criterion; an input to the trained evaluation model is image data of the user's non-verbal communication skills captured by the imaging means during conversation; The foreign language acquisition support system described in claim 1, characterized in that the output from the learned evaluation model is an evaluation of the user's non-verbal communication skills during conversation using the exemplary non-verbal communication skills as a judgment criterion by performing a predetermined calculation process.
7. 2. The foreign language acquisition support system according to claim 1, further comprising a curriculum creation means for creating a curriculum specialized for the user in order to compensate for the non-verbal communication skills of the user who has been evaluated poorly by the evaluation means.
8. 8. The foreign language acquisition assistance system according to claim 7, further comprising a learning material creation means for creating learning materials in accordance with the curriculum created by said curriculum creation means.
9. 9. The foreign language acquisition assistance system according to claim 8, wherein the learning materials are videos, and the videos are interactive videos in which the user and characters converse with each other.
10. The device further includes a subtitle generating means for displaying subtitles on the screen of the moving image, The foreign language acquisition assistance system according to claim 9, wherein, when a non-verbal communication skill of the user that has been poorly evaluated by the evaluation means appears during the dialogue, the subtitle generation means displays at least one of a subtitle indicating that fact and advice regarding the non-verbal communication skill on the screen.
11. The apparatus further includes a subtitle audio conversion means for converting the subtitles into audio, 11. The foreign language acquisition assistance system according to claim 10, wherein the subtitle audio conversion means converts the subtitles into audio and outputs the audio from the screen instead of or together with the subtitles.
12. A portable wireless communication device incorporating the foreign language learning assistance system according to any one of claims 1 to 11.
13. A foreign language acquisition support method for supporting a user in acquiring a foreign language, comprising: a first step of capturing an image of the user making a sound in accordance with the audio of the video; a second process of evaluating the user's nonverbal communication skills by comparing the user's nonverbal communication skills captured in the first process with exemplary nonverbal communication skills stored in a storage means that stores in advance exemplary facial expressions, gestures, and other exemplary nonverbal communication skills of the interlocutors during conversation, using a trained evaluation model that evaluates the interlocutors' nonverbal communication skills during conversation; a third step of displaying the evaluation created in the second step; A foreign language acquisition support method that includes:
14. The foreign language acquisition assistance method of claim 13, wherein in the second step, the user's non-verbal communication skills are evaluated based on at least one evaluation item of facial expressions, gaze, gestures, body language, proxemics, physical appearance, visual focus, auditory information, and cultural background.
15. The foreign language acquisition support method described in claim 13, characterized in that in the second step, feature quantities that quantitatively indicate the characteristics of each of the evaluation items are extracted from the user's image data, and these feature quantities are compared with exemplary non-verbal communication skills stored in the storage means.
16. The foreign language acquisition support method described in claim 13, characterized in that in the second process, the content of the user's conversation is used as auxiliary language data when evaluating the user's non-verbal communication skills.
17. The foreign language acquisition assistance method according to claim 13, characterized in that in the second step, cultural background information of the user is read from a cultural background information database that stores cultural background information of each country or region, and an adjustment is made to the evaluation of the user's non-verbal communication skills taking into account the cultural background information of the user.
18. the trained evaluation model is generated by machine learning to evaluate the non-verbal communication skills of the interlocutor using exemplary non-verbal communication skills stored in the storage means as a criterion; The input to the trained evaluation model is image data capturing the user's non-verbal communication skills during conversation, The foreign language acquisition assistance method described in claim 13, characterized in that the output from the learned evaluation model is an evaluation of the user's non-verbal communication skills during conversation using the exemplary non-verbal communication skills as a judgment criterion by performing a predetermined calculation process.
19. The foreign language acquisition support method according to claim 13, further comprising a fourth step of creating a curriculum specialized for the user to compensate for the user's non-verbal communication skills that were evaluated poorly in the second step.
20. 20. The foreign language acquisition assistance method according to claim 19, further comprising a fifth step of creating learning materials in accordance with the curriculum created in the fourth step.
21. the learning materials are videos, 21. The foreign language acquisition assistance method according to claim 20, wherein the video is an interactive video in which the user and characters converse with each other.
22. 22. The foreign language acquisition assistance method according to claim 21, further comprising a fifth step of, if, during the dialogue, a non-verbal communication skill of the user that was poorly evaluated in the second step appears, displaying at least a subtitle to that effect and advice regarding the non-verbal communication skill on the screen.
23. a sixth step of converting the subtitles into audio; a seventh step of playing a voice of the subtitles from the screen instead of or together with the subtitles; 23. The method of claim 22, further comprising:
24. A program for causing a computer to execute the foreign language learning assistance method according to any one of claims 13 to 23.
25. A portable wireless communication device having the program according to claim 24 stored therein.
Citation Information
Patent Citations
Language instruction system
JP2001282097A
Information processing device
JP2022142158A
Public Speaking Trainer With 3-D Simulation and Real-Time Feedback
US20190392730A1
Automated speech coaching systems and methods
US20210201696A1
KR20240029491A