Information processing method, information processing system, and computer program

The method determines and visually presents learning level differences using machine learning, addressing the lack of comprehensive feedback in existing technologies to enhance user learning efficiency.

JP7782468B2Active Publication Date: 2025-12-09SONY GROUP CORP
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2022576997
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-21
Filing Date
2021-11-16
Publication Date
2025-12-09
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

Existing learning support technologies fail to provide comprehensive feedback to users, as simply knowing their learning progress does not help identify future challenges or recognize differences from target learning objectives.

Method used

An information processing method that determines a user's learning level based on time-series media information using self-trained machine learning models, incorporating an attention mechanism to identify insufficient actions or behaviors, and visually presents these differences for improvement.

Benefits of technology

Enables users to understand and efficiently address areas of improvement in their learning by visually highlighting and quantifying the differences between their actions and desired outcomes, even without a tutor or professional guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007782468000001
    Figure 0007782468000001
  • Figure 0007782468000002
    Figure 0007782468000002
  • Figure 0007782468000003
    Figure 0007782468000003
Patent Text Reader

Abstract

Provided is an information processing method for performing processing for determining a learning level of a user or presenting the determined learning level. This information processing method involves: an input step for inputting time-series medium information representing the motion or action of a user during learning; a first determination step for determining the learning level of the user on the basis of the time-series medium information; and an output step for outputting, on the basis of the learning level of the user determined in the first determination step, a part where the motion or action of the user is different from a reference motion or action on the time-series medium information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to an information processing method, an information processing system, an information terminal, and a computer program that perform processing to support a user's learning. [Background technology]

[0002] In recent years, information technology has come to be used to support users' learning in language learning, musical instrument learning, sports (golf, baseball, soccer, etc.), and other training. For example, a speech learning system has been proposed that, when learning a second language through audio, uses a level determination program executed by a computer to determine the learning level from audio data based on the learner's speech and adjusts the playback speed of second language sentences to suit the learner's level (see Patent Document 1). Also proposed is an information processing device that acquires sensor information indicating information related to a first user playing golf from a sensor attached to a golf club, and acquires feedback information on first generated information based on the sensor information from a terminal of a second user and transmits it to the terminal of the first user (see Patent Document 2). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-113904 [Patent Document 2] WO2018 / 220948 Summary of the Invention [Problem to be solved by the invention]

[0004] An object of the present disclosure is to provide an information processing method, an information processing system, an information terminal, and a computer program that perform processing to support a user's learning. [Means for solving the problem]

[0005] The present disclosure has been made in consideration of the above problems, and a first aspect thereof is: an input step of inputting time-series media information representing the user's actions or behavior during learning; a first determination step of determining a learning level of a user based on time-series media information; an output step of outputting a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior based on the learning level of the user determined in the first determination step; The information processing method includes: wherein the first determination step processes time-series media information using a self-trained first machine learning model, and then determines the user's learning level using a supervised second machine learning model; and the first determination step uses an attention mechanism incorporated in the second machine learning model to determine, based on the time-series media information, the basis for determining that the user's learning level is insufficient or that the user needs to learn.

[0006] In the output step, a portion of the time-based media information where the user's action or behavior differs from a reference action or behavior is output to a presentation device. The information processing method according to the first aspect further includes a first presentation step of presenting to the user through the presentation device the portion of the time-based media information where the user's action or behavior differs from a reference action or behavior. In the first presentation step, the portion of the time-based media information where the user's action or behavior differs from a reference action or behavior is visually presented.

[0007] The information processing method according to the first aspect further includes a second determination step of determining distance information representing a difference between the user's movement or action and a reference movement or action, and a second presentation step of outputting the determination result of the second determination step to a presentation device and presenting it to the user. In the second presentation step, the distance information is visually presented in an N-dimensional space with the reference movement or action positioned at the center.

[0008] Furthermore, a second aspect of the present disclosure is an input unit for inputting time-series media information representing the actions or behavior of a user during learning; a first determination unit that determines a learning level of a user based on time-series media information; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior based on the learning level of the user determined by the first determination unit; The information processing system according to the second aspect may further include a sensor unit that detects a movement or behavior of a user during learning to acquire time-series media information, and a presentation device that outputs a portion of the time-series media information in which the movement or behavior of the user differs from a reference movement or behavior.

[0009] However, the term "system" used here refers to a logical collection of multiple devices (or functional modules that realize specific functions), regardless of whether each device or functional module is contained within a single housing. In other words, both a single device consisting of multiple parts or functional modules and a collection of multiple devices are considered "systems."

[0010] Furthermore, a third aspect of the present disclosure is a sensor unit that detects the actions or behavior of a user during learning and acquires time-series media information; a communication unit that transmits the time-series media information to an external device and receives from the external device a determination result of the user's learning level and a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior; a presentation unit that presents the received information; It is an information terminal equipped with the above.

[0011] Furthermore, a fourth aspect of the present disclosure is an input unit for inputting time-series media information representing the actions or behavior of a user during learning; a first determination unit that determines a learning level of a user based on time-series media information; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior, based on the learning level of the user determined by the first determination unit; It is a computer program written in a computer-readable format so that the computer functions as

[0012] A computer program according to a fourth aspect of the present disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. In other words, by installing the computer program according to the fourth aspect of the present disclosure on a computer, a cooperative action is exerted on the computer, and the same effects as those of the information processing method according to the first aspect of the present disclosure can be obtained. [Effects of the Invention]

[0013] According to the present disclosure, it is possible to provide an information processing method, an information processing system, an information terminal, and a computer program that perform processing to determine a user's learning level or to present the determined learning level.

[0014] It should be noted that the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited to these. Furthermore, the present disclosure may also bring about additional effects in addition to the effects described above.

[0015] Further objects, features, and advantages of the present disclosure will become apparent from the following detailed description based on the embodiments and accompanying drawings. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 is a diagram showing the basic configuration of an information processing system 100 that supports a user's learning. [Figure 2] FIG. 2 is a flowchart showing an example of the operation of the information processing system 100. [Figure 3]FIG. 3 is a flowchart showing another example of the operation of the information processing system 100. [Figure 4] FIG. 4 is a diagram showing an example of visually presenting the determination result. [Figure 5] FIG. 5 is a diagram showing an example of visually presenting the determination result. [Figure 6] FIG. 6 is a diagram showing an example of visually presenting the determination result. [Figure 7] FIG. 7 is a diagram showing an example of a system configuration. [Figure 8] FIG. 8 is a diagram showing another example of a system configuration. [Figure 9] FIG. 9 is a diagram showing yet another example of a system configuration. [Figure 10] FIG. 10 is a diagram showing yet another example of a system configuration. [Figure 11] FIG. 11 is a diagram showing an example in which distance information indicating the difference between a user's movement or action and a reference movement or action is presented on a two-dimensional plane. [Figure 12] FIG. 12 is a diagram showing how distance information changes as learning progresses. [Figure 13] FIG. 13 is a diagram showing an example in which distance information indicating the difference between a user's movement or action and a reference movement or action is presented in a three-dimensional space. [Figure 14] FIG. 14 is a diagram showing an example of the internal configuration of the determination unit 103 made up of a DNN. [Figure 15] FIG. 15 is a diagram for explaining the operation of the attention mechanism. [Figure 16] FIG. 16 is a diagram for explaining a method for calculating the Triplet Loss. [Figure 17] FIG. 17 is a diagram showing an example of the internal configuration of the determination unit 103 that focuses on differences in learning methods. [Figure 18] FIG. 18 is a diagram for explaining the self-learning method of the self-learning model 1701. As shown in FIG. [Figure 19] FIG. 19 is a diagram showing the relationship between the content of learning that can be supported by the present disclosure and the type of time-series media information. [Figure 20] FIG. 20 is a diagram showing an example of the configuration of a UI that presents the user with the results of the determination of the user's learning level. [Figure 21] FIG. 21 is a diagram showing an example of the configuration of a UI that presents the user with the results of the determination of the user's learning level. [Figure 22] FIG. 22 is a diagram showing an example of the configuration of a UI that presents the user with the results of the determination of the user's learning level. [Figure 23] FIG. 23 is a diagram showing an example of the configuration of an information processing device 2300. [Figure 24] FIG. 24 is a diagram showing an example of the configuration of an information terminal 2400. DETAILED DESCRIPTION OF THE INVENTION

[0017] The present disclosure will be described below in the following order with reference to the drawings.

[0018] A. Overview B.Basic configuration B-1. Functional Blocks B-2. System Operation B-3. ​​How to present the results B-4. Specific system configuration examples B-5. Displaying distance information C. Implementation using machine learning models C-1. Configuring machine learning models C-2. Learning Methods D. Application Examples E.UI Example F. Equipment configuration example F-1. Example of information processing device configuration F-2. Example of information terminal configuration

[0019] A. Overview In recent years, information technology has come to be used to support users' learning in areas such as language learning, learning a musical instrument, and training for sports (golf, baseball, soccer, etc.). For example, a learner's level can be determined using a computer (see Patent Document 1). However, simply presenting a level indicating the user's learning progress is not sufficient feedback to the user. In other words, simply knowing one's own level makes it difficult for the user to identify future challenges, and the user cannot recognize where or how the current level differs from the target learning objectives.

[0020] Therefore, this disclosure proposes a method for determining the learning level of a user's actions and behaviors based on time-series media information such as video and audio that represent the actions and behaviors performed by the user, presenting the determination result to the user, and if the determination result indicates that learning is still insufficient, further presenting learning progress information such as which parts of the user's actions and behaviors are insufficient and to what extent, as well as a device for realizing this method.

[0021] In this specification, unless otherwise specified, a user refers to a "learner" who strives to learn a language, a musical instrument, a sport, or the like.

[0022] For example, when learning a second language, a user studies to make their spoken voice and written sentences closer to those of a native speaker. According to a method disclosed herein, it is possible to present to the user the results of determining whether the user's spoken voice and written sentences are similar to those of a native speaker. Furthermore, according to the present disclosure, it is possible to visually present the extent to which the user's spoken voice and written sentences differ from those of a native speaker. Therefore, even when learning a second language without a native speaker or tutor (i.e., studying alone), a user can understand the differences between their current pronunciation and writing and the pronunciation and writing they should aim for, and can efficiently train for language acquisition. Of course, the present disclosure can be applied not only to language learning but also to learning various user actions and behaviors that produce sound, such as singing, playing an instrument, giving a speech, acting, and stand-up comedy.

[0023] Furthermore, when a user learns a sport (ball games such as golf, tennis, soccer, and baseball, or martial arts such as judo, karate, kendo, and boxing), the user trains to make their physical movements (swings, running kicks, techniques, defensive moves, etc.) more similar to those of a professional athlete or trainer. According to a method disclosed herein, a determination of whether the user's physical movements are similar to those of a professional athlete or trainer is made based on a video of the user training or playing, and the results are presented to the user. The method also visually indicates which parts and to what extent the user's physical movements differ from those of a professional athlete or the trainer's instructions. Therefore, even if a professional athlete or trainer is not nearby, a user can understand the difference between their current physical movements and the physical movements they should aim for when training for a sport, and can train efficiently. Of course, the present disclosure can be applied not only to sports, but also to various physical movements the user learns, such as playing a musical instrument, calligraphy, cooking, public speaking, acting, stand-up comedy, and comedy sketches.

[0024] B.Basic configuration B-1. Functional Blocks FIG. 1 schematically shows the basic configuration of an information processing system 100 that supports user learning by applying the present disclosure.

[0025] The sensor unit 101 includes an image sensor such as a camera and an audio sensor such as a microphone that detect video and audio representing the user's actions and behaviors. The sensor unit 101 outputs time-series media information such as video and audio representing the user's actions and behaviors.

[0026] The determination unit 103 receives time-series media information, such as video and audio, representing the user's actions and behaviors from the sensor unit 103 via the input unit 102. The determination unit 103 then determines the user's learning level based on the time-series media information and presents the determination result to the user. For example, when a user is learning a second language, the determination unit 103 determines whether the user's speech sounds similar to that of a native speaker, i.e., whether the user's pronunciation is comparable to that of a native speaker and no longer needs to be studied, or whether the user's pronunciation is different from that of a native speaker and therefore needs to continue studying, and outputs a determination result on whether or not the user needs to continue studying. Furthermore, when the determination unit 103 determines that the user needs to continue studying, it determines and outputs the extent to which the user's learning is insufficient in the time-series media information. The determination unit 103 performs a process of determining the user's learning level using a trained machine learning model; details of this point will be described later.

[0027] The presentation unit 104 presents to the user the result of the determination of whether the user needs to study, output from the determination unit 103, and the portion of the time-series media information that is determined to be insufficient for study (however, if it is determined that the user needs to continue studying). The presentation unit 104 is equipped with a display that visually presents the determination result by the determination unit 103, but may also be equipped with an audio output device such as a speaker so that the result can also be presented by audio announcement. In particular, if the relevant portion of the time-series media information is visually presented on the display screen, it is easy for the user to understand which part of their actions or behaviors is insufficient and to what extent.

[0028] B-2. System Operation 2 shows in the form of a flowchart an example of an operation of the information processing system 100. This operation is triggered, for example, by an instruction to determine the learning level from a user who is currently learning.

[0029] The sensor unit 101 detects video and audio representing the actions and behaviors of the user during learning using an image sensor and an audio sensor, and outputs the video and audio as time-series media information (step S201).

[0030] The determination unit 103 inputs the time-series media information via the input unit 102 and determines the learning level of the user's movements and actions (step S202). Also, in step S202, if there is a portion where learning is insufficient, the determination unit 103 determines which portion of the time-series media information is insufficient and to what extent. The determination unit 103 performs a process of determining the user's learning level from the time-series media information using a trained machine learning model. Then, if the user's movements and actions are close to the reference movements and actions, the determination unit 103 determines that the user's learning is sufficient (Yes in step S203), but if the user's movements and actions are not close to the reference movements and actions, the determination unit 103 determines that the user's learning is insufficient (No in step S203).

[0031] For example, when the present disclosure is applied to learning a second language, the determination unit 103 determines that the user's learning is sufficient if the user's voice or sentences are at the level of a native speaker, but otherwise determines that the user's learning is insufficient.Furthermore, when the present disclosure is applied to sports training, the determination unit 103 determines that the user's learning is insufficient if the user's body movements are similar to those of a professional athlete or trainer.

[0032] If the determination unit 103 determines that the user's learning is sufficient (Yes in step S203), the determination result that learning is sufficient or learning has ended is presented to the user via the presentation unit 104 (step S204), and this process ends. Furthermore, even if the determination unit 103 determines that the user's learning is sufficient, if there are parts where learning is insufficient, the determination unit 103 may present the parts of the time-series media information where learning is insufficient, or may allow the user to continue learning if they wish.

[0033] On the other hand, if the determination unit 103 determines that the user's learning is insufficient (No in step S203), the determination result that the learning is insufficient or that the learning should be continued is presented to the user via the presentation unit 104 (step S205), and the part of the time-series media information where learning is insufficient is presented (step S206), and this process ends.

[0034] 3 is a flowchart showing another example of the operation of the information processing system 100. This operation is triggered, for example, by an instruction to determine the learning level from a user who is currently learning.

[0035] The sensor unit 101 detects video and audio representing the user's movements and actions during learning using an image sensor and an audio sensor, and outputs the video and audio as time-series media information (step S301). The determination unit 103 then inputs the time-series media information via the input unit 102 and determines the learning level of the user's movements and actions (step S302). In addition, in step S302, if there is a part where learning is insufficient, the determination unit 103 determines where in the time-series media information the learning is insufficient and to what extent.

[0036] Here, if the determination unit 103 determines that the user's learning is sufficient (Yes in step S303), the determination result that the learning is sufficient or that the learning has ended is presented to the user via the presentation unit 104 (step S304), and this process ends. Furthermore, even if the determination unit 103 determines that the user's learning is sufficient, if there are parts where the learning is insufficient, the determination unit 103 may present the parts of the time-series media information where the learning is insufficient, or may allow the user to continue learning if they wish.

[0037] On the other hand, if the determination unit 103 determines that the user's learning is insufficient (No in step S303), it presents the determination result that learning is insufficient or that learning should be continued to the user through the presentation unit 104 (step S305), and also presents the part of the time-series media information where learning is insufficient (step S306). Thereafter, the process returns to step S301, and the user's learning and the detection and determination of time-series media information that represents the user's actions and behaviors during learning are repeatedly performed until it is determined that the user's learning is sufficient (Yes in step S303).

[0038] B-3. ​​How to present the results 4 shows an example in which the presentation unit 104 visually presents which parts of the time-series media information are insufficiently learned and to what extent. Here, an example is assumed in which, when a user is learning English pronunciation, a speech waveform signal when the user utters the phrase "This was easy for us" is input to the system 100 as time-series media information. Note that the phrase "This was easy for us" may be presented on the screen by an English learning program and read aloud by the user, or may be freely uttered by the user.

[0039] When the determination unit 103 determines that the phrase "This was easy for us" spoken by the user is different from that of a native speaker and that the user has insufficient learning in some areas, the determination unit 103 determines which parts of the speech waveform signal are insufficient in learning and to what extent. Then, the presentation unit 104 highlights the parts of the speech waveform signal that are different from that of a native speaker, as shown in FIG. 4. Furthermore, the presentation unit 104 displays the phrase "This was easy for us" spoken by the user in text along with the speech waveform signal, as shown in FIG. 4, and highlights the words or character strings "This," "eas," and "for" whose pronunciation is determined to be different from that of a native speaker. Note that the method of highlighting is not particularly limited. For example, methods other than (or in addition to) displaying the relevant words or character strings in high brightness include increasing the character size, making them bold, changing the font, and circling them.

[0040] Therefore, the user not only recognizes that their pronunciation is different from that of a native speaker, but also easily understands the parts (words or character strings) that are pronounced differently from that of a native speaker.The user can then efficiently learn the language by focusing on the parts that have been pointed out, such as paying particular attention to the pointed out parts "This," "eas," and "for" and correcting the pronunciation of the same phrase.

[0041] As in the example of operation shown in FIG. 3, examples of visual presentations in which detection and determination of time-series media information are repeatedly performed until it is determined that the user's learning is sufficient are shown in FIGS.

[0042] Based on the visual presentation as shown in Fig. 4, the user pronounces the phrase "This was easy for us", paying particular attention to the words or character strings "This", "eas", and "for" that were identified as being different from those of a native speaker. As a result, the user's pronunciation improves, and the determination unit 103 determines that only the word "This" differs from that of a native speaker. In this case, the presentation unit 104 highlights the part corresponding to the word "This" in the input speech waveform signal, and highlights only the word "This" in the text display of the phrase "This was easy for us", as shown in Fig. 5.

[0043] Therefore, the user can understand that his / her pronunciation is closer to that of a native speaker than the previous time, and that he / she should pay particular attention to pronouncing the word "This" next time. As a result, the user's pronunciation is further improved, and he / she can pronounce the entire phrase "This was easy for us" at a level comparable to that of a native speaker. As a result, the determination unit 103 determines that learning is sufficient, and can end learning the pronunciation of the phrase "This was easy for us". Furthermore, as shown in FIG. 6, the presentation unit 104 no longer presents parts of the input speech waveform signal and the text display of the phrase "This was easy for us" that differ from those of a native speaker.

[0044] Although not shown in Figures 4 to 6, not only can the portions of the time-series media information where learning is insufficient (e.g., portions of the audio waveform signal that differ from that of a native speaker) be highlighted, but the degree of insufficient learning or the basis for determining that learning is insufficient can also be quantified and presented. Furthermore, the highlighting level can be adjusted according to the numerical value of the basis (e.g., the higher the numerical value, the higher the brightness or the larger the font size). Also, specific instructions on how to improve, such as "You're pronouncing 'this' as 'zis'. Try pronouncing the 'th' sound by placing the tip of your tongue between your upper and lower teeth," can be displayed as a pop-up on the screen, a video demonstrating how to move your mouth, or audio guidance can be output. Visually representing the level of learning in this way can help users better understand the areas they need to focus on studying.

[0045] B-4. Specific system configuration examples 7 shows an example of a system configuration in which the sensor unit 101, the input unit 102, the determination unit 103, and the presentation unit 104 are all mounted on a single device 700. The single device 700 referred to here may be, for example, a multi-function information terminal such as a smartphone or tablet owned by a user, a personal computer, or a device manufactured specifically for learning support. However, some of the components such as the sensor unit 101 and the presentation unit 104 may be configured to be externally connected to the device 700 rather than being built into the device 500. For example, some of the components may be externally connected to the device 500 using a wired interface such as USB (Universal Serial Bus) or HDMI (High Definition Multimedia Interface), or a wireless interface such as Bluetooth (registered trademark) or Wi-Fi (registered trademark).

[0046] FIG. 8 shows an example of a system configuration separated into a first device 801 equipped with a sensor unit 101 and a second device 802 including an input unit 102, a determination unit 103, and a presentation unit 104. The first device 801 and the second device 802 are interconnected via a wireless or wired interface. The first device 801 includes a camera, a microphone, and the like installed in a location where a user's movements and actions during learning can be easily detected. The first device 801 may be a sensor attached to a tool used by the user in a sports game, such as a swing sensor attached to a golf club. On the other hand, the second device 802 may be, for example, a multi-function information terminal such as a smartphone or tablet owned by the user, or a personal computer. The second device 802 determines the user's learning level and the extent to which the user's movements and actions are insufficient based on time-series information acquired from the first device 801 via a wireless or wired interface, and visually presents the results.

[0047] FIG. 9 shows an example of a system configuration separated into a first device 901 equipped with a sensor unit 101, a second device 902 equipped with an input unit 102 and a judgment unit 103, and a third device 903 equipped with a presentation unit 104. The first device 901 includes a camera and a microphone installed in a location where the user's movements and actions during learning can be easily detected. The first device 901 may also be a sensor attached to a tool used by the user in a sports competition, such as a swing sensor attached to a golf club. The second device 902 includes a device with high computing power, such as a personal computer or a cloud computer. The third device 903 includes a multi-functional information terminal, such as a smartphone or tablet, owned by the user, and mainly performs the process of receiving the judgment results from the second device 902 and presenting them to the user. When the second device 902 is a cloud computer, the system can also be configured to provide a learning level judgment processing service to multiple or many users.

[0048] FIG. 10 shows an example of a system configuration separated into a first device 1001 equipped with a sensor unit 101 and a presentation unit 104, and a second device 1002 equipped with an input unit 102 and a determination unit 103. The first device 1001 may be, for example, a multi-function information terminal such as a smartphone or tablet carried by a user, or a personal computer. The sensor unit 101 may not be built into the first device 1001 but may be externally connected to the first device 1001. On the other hand, the second device 1002 may be, for example, a cloud computer. The second device 1002 receives time-series media information from the first device 1001 and returns a determination result of the user's learning level based on the time-series media information to the first device 1001. Thus, the first device 1001 acquires time-series media information from a user who is learning and transmits (uploads) it to the second device 1002, and then receives (downloads) the determination result from the second device 1002 and presents it to the user. When the second device 1002 is a cloud computer, the system can be configured to provide a learning level determination processing service for multiple or multiple users.

[0049] B-5. Displaying distance information In the explanation so far, the determination unit 103 determines the learning level of the user based on time-series media information such as video and audio that represent the actions and behaviors performed by the user, and if there is a part where the learning level is insufficient, determines which part of the time-series media information is insufficient and to what extent, and presents this to the user via the presentation unit 104. As an advanced version of this, the determination unit 103 may determine how much the entire actions and behaviors performed by the user differ from a reference action or behavior (specifically, an ideal action or behavior of a native speaker, professional athlete, or trainer) as N-dimensional (two-dimensional or three-dimensional) distance information, and present this via the presentation unit 104. The determination unit 103 can determine the distance information using a distance learning model, but details of this point will be described later.

[0050] 11 shows an example in which the difference between a user's motion or action and a reference motion or action is expressed as distance information on a two-dimensional plane and presented by the presentation unit 104. In the example shown in FIG. 11, a graphic 1101 representing the reference motion or action is displayed at the center of a two-dimensional plane 1100, and a graphic 1102 representing the user's motion or action is displayed at a distance from the center. In the case where a user is learning the pronunciation of a second language, the central graphic 1101 represents the pronunciation of a native speaker, the peripheral graphic 1102 represents the user's current pronunciation, and the distance from the center to the position where graphic 1102 is placed represents the user's current pronunciation level. In the case where a user is training their baseball bat swing, the central graphic 1101 represents an ideal (or professional baseball player's) bat swing, the peripheral graphic 1102 represents the user's current bat swing, and the distance from the center to the position where graphic 1102 is placed represents the user's current bat swing. In the case of training a golf swing, the central figure 1101 represents an ideal (or professional golfer's) golf swing, the surrounding figure 1102 represents the user's current golf swing, and the distance from the center to the position where figure 1102 is placed represents the user's current golf skill.

[0051] FIG. 12 shows how distance information changes on a two-dimensional plane 1100 as the user's learning progresses. In the early stages of learning, as shown in FIG. 12(A), a figure 1102 indicating the user's current learning level is far away from a central figure 901 on the two-dimensional plane 900. Thereafter, as the user continues training (e.g., correcting their pronunciation), the figure 1102 indicating the user's current learning level gradually approaches the central figure 1101, as shown in FIGS. 12(B) to 12(D). During the training process, the orientation of the figure 1102 changes from the central figure 1101, which indicates a phenomenon such as the movement of a part that is different from the reference action or behavior in the time-series media information. Then, as the training continues, the figure 1102 changes its orientation and moves closer to the central figure 1101, and eventually, as shown in FIG. 12(E), the figure 1102 overlaps with the figure 1101. This indicates that the difference between the user's actions and behaviors and the reference actions and behaviors has become small enough that learning has ended.

[0052] Based on the visual presentation of distance information as shown in Figure 11, the user can understand whether their own movements or actions are close to or far from the reference movements or actions. In addition, by visually observing the changes in distance information during the training process as shown in Figure 12, the user can determine for themselves, without the help of a trainer, whether their movements or actions are improving as a result of training, i.e., whether the direction of their training is correct.

[0053] Furthermore, when a user determines that their own movements or actions are far from the reference movements or actions based on the visual presentation of distance information as shown in Figure 11, they can use the visual presentation as shown in Figures 4 to 6 to check the parts of the time-series media information that differ from the reference movements or actions, understand why their movements or actions are different from the ideal, and use this information as a reference for future training methods.

[0054] 11 and 12 show examples of distance information represented on a two-dimensional plane 1100, but distance information may also be represented in a three-dimensional space as shown in Fig. 13. In the example shown in Fig. 13, a figure 1301 indicating a reference movement or action is displayed in the center of a three-dimensional space 1300, and a figure 1302 indicating a user's movement or action is displayed at a distance from the center. Fig. 12 shows how the distance information changes on the two-dimensional plane 1100 as the user progresses in their studies, but by using a three-dimensional space 1300 as shown in Fig. 13, the change in distance information as the user progresses in their studies can be shown in a more expressive manner.

[0055] C. Implementation using machine learning models C-1. Configuring machine learning models In the above section B, it has been explained that the determination unit 103 has a function of determining the learning level of a user based on time-series media information that represents the actions and behaviors performed by the user during learning, and of determining where the actions and behaviors performed by the user in the time-series media information differ from the reference actions and behaviors. Such a function in the determination unit 103 can be realized using a trained machine learning model such as a DNN (Deep Neural Network).

[0056] A classification model using machine learning such as a DNN generally comprises a feature extraction unit that extracts features from input data such as time-series media information, and a classification unit that classifies output labels based on the extracted features. In this embodiment, the classification unit classifies the label as whether the user has learned or not. Specifically, the feature extraction unit comprises, for example, a convolutional neural network (CNN), and the classification unit comprises a fully connected layer. Furthermore, by incorporating an attention mechanism into the classification unit, it is possible to indicate in the time-series media information the basis for labeling a user as not having learned enough. Note that attention is one of the implementation methods of XAI (eXplainable AI) technology, which explains the basis for the judgments of a machine learning model, and is well known in the industry as a method of incorporating a mechanism that indicates points of interest in input data (i.e., an attention mechanism) into a machine learning model.

[0057] 14 is a schematic diagram showing an example of the internal configuration of the determination unit 103 made up of a DNN 1400. The internal configuration of the illustrated determination unit 103 will be described below.

[0058] The feature extraction unit 1410 comprises a plurality of (T in the illustrated example) CNNs 1411-1, 1411-2, ..., 1411-T. Each of the CNNs 1411-1, 1411-2, ..., 1411-T extracts time-series media information (e.g., a voice waveform signal pronounced by a user) 1401 over a predetermined time interval P1, P2, ..., P3, ..., P4. T , 1411-T, a feature value z1 is calculated by summarizing the outputs of a predetermined number of CNNs 1411-1, 1411-2, ..., 1411-T that have received a plurality of consecutive interval data. ´ , z2 ´, …, z T ´ and outputs it to the subsequent classification unit 1420 or a distance learning model (described later) that estimates distance information.

[0059] As will be described later, a self-learning model is used in the feature extraction unit 1410. This is used not only to facilitate classification in the subsequent classification unit 1420, but also to overcome the difficulty of collecting data on user movements and actions. When handling audio signals as time-series media information, such as when training a user's pronunciation in language learning, the feature extraction unit 1410 can use, for example, wav2vec or wav2vec2.0. When handling video as time-series media information to train a user's physical movements in sports, the feature extraction unit 1410 can use CVRL (Contrastive Video Representation Learning) or PCL (Pretext-Contrastive Learning).

[0060] The classification unit 1420 classifies whether the user's actions and behaviors are close to the reference actions and behaviors, i.e., whether the user has sufficiently learned, based on the feature amounts of the user's actions and behaviors extracted by the feature extraction unit 1410. In the example shown in Fig. 14, the classification unit 1420 includes a Bi-LSTM layer 1421 and a classifier layer 1423, and also has an attention mechanism 1422 built in.

[0061] The Bi-LSTM (Bidirectional LSTM) layer 1421 is a neural network called a bidirectional LSTM (Long Short-Term Memory), which is an improved version of an RNN (Recurrent Neural Network), and integrates the results of running LSTM from left to right and the results of running LSTM from right to left.

[0062] Here, we briefly explain neural networks. A neural network consists of three layers: an input layer, a hidden layer (or a hidden layer), and an output layer. Each layer contains a required number of unit elements called neurons. The neurons in the input layer and the hidden layer are connected with weights. Similarly, the neurons in the hidden layer and the output layer are also connected with weights. The neural network learns by inputting data such as features and updating the weight coefficients to output correct recognition results using, for example, backpropagation, which backpropagates an error signal. The error signal here represents the difference between the output signal from the output layer and a teacher signal. An RNN is a neural network with an internal closed loop that can dynamically update its internal state while memorizing past information. An LSTM is constructed by replacing the hidden layer of an RNN with an LSTM block to model long-term data. An LSTM block has three gates: an input gate, a forget gate, and an output gate, as well as a memory cell. The memory cell represents the internal state and can retain long-term information. The input gate and the output gate respectively function to adjust the input data and the output data, and the forget gate functions to adjust the memory cell input from the previous time.

[0063] Referring again to FIG. 14, the description of the DNN 1400 will continue. The classifier layer 1423 determines whether the user's learning is sufficient based on the inference result of the time-series media information by the Bi-LSTM layer 1421, and outputs the result. For example, in the case of language learning, the classifier layer 1423 determines whether the user's pronunciation is at the level of a native speaker. In addition, in the case of sports training, the classifier layer 1423 determines whether the user's physical movements are at the level of a professional athlete or trainer.

[0064] When the classifier layer 1423 determines that the user's learning is insufficient, the attention mechanism 1422 detects portions of the input time-series media information that support the determination that the user's learning is insufficient. Attention is one implementation method of XAI technology and is well known in the industry as a technique for incorporating a mechanism that indicates points of focus for input data into a machine learning model when the model classifies. For example, in the case of language learning, when the classifier layer 1423 determines that the user's pronunciation is different from that of a native speaker, the attention mechanism 1422 can identify portions of the speech waveform signal that differ from that of a native speaker, as shown in FIG. 4. Furthermore, when a correspondence is established between the speech waveform signal and a character string (e.g., "This was easy for us"), the attention mechanism 1422 can also identify words or character strings determined to be different from those of a native speaker.

[0065] The operation of the attention mechanism 1422 will be described with reference to Fig. 15. However, in this figure, it is assumed that a raw audio waveform signal 1501 is input as time-series media information. The feature extraction unit 1410, which is configured with wav2vec or the like, extracts a feature quantity z1 from the audio waveform signal 1501. ´ , z2 ´ ,z3 ´ , …, z T ´ Then, the classification unit 1423 (not shown in FIG. 15) outputs the feature amount z1 ´ , z2 ´ ,z3 ´ , …, z T ´Based on this, the attention mechanism 1422 determines that the user's speech waveform signal 1501 is different from the reference (or ideal) speech waveform signal of a native speaker. At this time, the attention mechanism 1422 indicates which of the time intervals of the input speech waveform signal the classification unit 1423 focused on when making the determination. In the example shown in FIG. 15, the attention mechanism 1422 expresses the degree to which each time interval of the speech waveform signal contributed to the determination result of "different from the ideal speech waveform signal" as a numerical value between 0 and 1, 0.38, 0.71, 0.42, ..., 0.92. Note that a higher numerical value indicates a greater contribution to the determination result. Based on the output of the attention mechanism 1422, the presentation unit 104 can highlight the portions of the speech waveform signal that differ from those of the native speaker, as indicated by reference numeral 1502. Furthermore, the strength of the highlighting of the multiple portions that differ from those of the native speaker is changed according to the calculated numerical value.

[0066] 14, as mentioned in the above section B-5, the determination unit 103 determines distance information that indicates how much the entire action or behavior performed by the user differs from a reference action or behavior using a distance learning model. The distance learning model uses the feature quantity z1 extracted by the feature extraction unit 1410 from the time-series media information that is the input data. ´ , z2 ´ ,z3 ´ , …, z T ´ It is trained to estimate the distance between the user's actions and behaviors and the reference actions and behaviors from a feature vector with elements.

[0067] Basic loss functions such as Contrastive Loss and Triplet Loss can be used to train a distance learning model. Here, Contrastive Loss is a loss calculated based on the distance between two points. As shown in Figure 16, Triplet Loss is a loss calculated as a set of three: a reference Anchor feature, a Positive feature with the same label as the Anchor, and a Negative feature that is different from the Anchor. Then, the Anchor, Positive, and Negative features are arranged as vectors in space, and the distance between the Anchor and the Positive is expressed as d p , the distance between Anchor and Negative is d n For example, L triplet =[d p -d n +α]+ can be defined as Triplet Loss (where α is a hyperparameter that represents the margin). The top of Figure 16 shows how to calculate Triplet Loss when the professional (or ideal) is positive and the learner is negative. The bottom of Figure 16 shows how to calculate Triplet Loss when the learner is positive and the professional (or ideal) is negative.

[0068] The determination unit 103 can determine distance information between the user's movement or action and a reference movement or action from time-series media information obtained by sensing the user's movement or action using the trained distance learning model. Then, as shown in Fig. 11, the presentation unit 104 can visualize and present the difference between the user's movement or action and the reference movement or action as distance information on a two-dimensional plane. Therefore, the user can understand whether their own movement or action is close to or far from the reference movement or action from the visualized information as shown in Fig. 11.

[0069] C-2. Learning Methods Next, a learning method for the machine learning model used in the determination unit 103 will be described.

[0070] Deep learning requires a large amount of training data. When attempting to use supervised training with DNN1400, which determines user behavior and actions, it is necessary to collect a huge amount of user behavior data (audio, video, etc.) and to label each piece of data, which requires an excessively large amount of annotation work. Without sufficient training, problems such as unstable operation of DNN1400 and incorrect judgments can occur.

[0071] The supervised learning of the classification unit 1420 at the downstream of the DNN 1400 is performed using motion data from the user (or a beginner at the same level as the user) and ideal motion data from professionals. However, collecting user motion data is often difficult. For example, in the case of language learning, it is relatively easy to collect speech data from native speakers through various media such as television and radio broadcasts and online video streaming services, but it is difficult to collect speech data pronounced by scholars.

[0072] Therefore, in this disclosure, we propose that the feature extraction unit 1410 at the front of the DNN 1400 used in the judgment unit 103 performs self-learning using large amounts of data collected through broadcasting, distribution services, etc., and that the classification unit 1420 at the rear performs supervised learning.

[0073] 17 shows an example of the internal configuration of the determination unit 103 that focuses on differences in learning methods. In the example shown, the determination unit 103 is composed of a self-learning model 1701, a supervised classification model 1702, and a distance learning model 1703. In addition, an attention mechanism is incorporated into the supervised classification model 1702.

[0074] The self-learning model 1701 corresponds to the feature extraction unit 1410 in Fig. 14. The self-learning model 1701 self-learns to find a good representation of time-series media information such as audio and video that the user wants to learn. A good representation means a representation that is easy to classify in the subsequent supervised classification model 1702. When supporting language learning, for example, a self-learning model for audio such as wav2vec or wav2vec2.0 can be used, and when supporting training of physical movements such as sports, for example, a self-learning model for video such as CVRL or PCL can be used.

[0075] The self-learning model 1701 not only facilitates classification in the subsequent supervised classification model 1702, but also overcomes the difficulty of collecting data on user movements and actions (time-series media information). The subsequent supervised classification model 1702 is trained using user movement data and ideal movement data of professionals, etc., but it is often difficult to collect movement data of users (or beginners at a similar level to the user). On the other hand, ideal movement data of professionals, etc. (such as speech data of native speakers) can be collected in large quantities through television and radio broadcasts, video distribution services via the Internet, etc. By self-learning using a large amount of collectable data, the self-learning model 1701 can obtain expressions that are easy to classify in the subsequent supervised classification model 1702.

[0076] The supervised classification model 1702 corresponds to the classification unit 1420 in Fig. 14 and is composed of a DNN that supports time series, such as an RNN or an LSTM, which is an improved version of an RNN. The supervised classification model 1702 classifies the user's motion data and the reference motion data based on the representation obtained by the self-learning model 1701 in the previous stage.

[0077] The attention mechanism is incorporated into the supervised classification model 1702 to visualize which parts of the user's motion data the supervised classification model 1702 focused on when classifying (see, for example, FIGS. 4 to 6). When the supervised classification model 1702 classifies a user's motion as different from the ideal motion of a professional or the like, the focus corresponds to the parts of the user's motion or behavior that differ from the ideal motion or behavior of a professional or the like. The attention mechanism outputs a numerical value between 0 and 1 representing the degree of attention paid to each section of the time-series media information during classification. Then, by visualizing and showing parts with high numerical values ​​on the time-series media information or by displaying numerical values ​​for each section of the time-series media information (see FIG. 15), the user can easily understand which parts of their motion or behavior differ from the ideal motion or behavior and to what extent.

[0078] The distance learning model 1703 calculates distance information that indicates how much the user's overall actions and behaviors differ from a reference action and behavior, based on the representation obtained by the self-learning model 1701 in the previous stage. Basic loss functions such as Contrastive Loss and Triplet Loss can be used for learning the distance learning model 1703 (as described above). Then, as shown in FIG. 11, the presentation unit 104 can visualize and present the difference between the user's actions and behaviors and the reference action and behavior as distance information on a two-dimensional plane.

[0079] Next, the self-learning method of the self-learning model 1701 will be described with reference to FIG.

[0080] The self-learning model 1701 corresponds to the feature extraction unit 1410 in FIG. 14 and is composed of a CNN. When handling an audio signal as time-series media information, for example, wav2vec or wav2vec2.0 can be used as the self-learning model 1701. When handling video as time-series media information to train a user's physical movements, such as in sports, CVRL or PCL can be used as the self-learning model 1701. FIG. 18 shows an example of using wav2vec2.0, and is a speech recognition framework using a Transformer. The speech recognition framework consists of an encoder unit 1801 composed of a CNN that convolves an audio signal into a latent representation, and a Transformer unit 1802 that obtains a contextual representation from the latent representation.

[0081] Each CNN in the encoder unit 1801 convolves interval data obtained by dividing a speech waveform signal into time intervals, and outputs latent representations Z. The transformer unit 1802 inputs quantized representations Q of the latent representations Z for each time interval to obtain context representations C. Then, using the contrastive loss of the latent representation Z for each time interval and the context representation C as a loss function, the self-learning model 1701 (i.e., the entire speech recognition framework) is self-trained so that the context representation C for each time interval approximates the latent representation Z for the corresponding time interval, but moves away from the latent representations Z for other time intervals.

[0082] In the self-trained speech recognition framework, the CNN encoding unit 1801 is used as the self-trained model 1701. In the context of FIG. 14, the CNN encoding unit 1801 is used as the feature extraction unit 1410.

[0083] When training supervised classification model 1702, training of self-learning model 1701 is stopped, and supervised data (i.e., labeled time-series media information) is input to self-learning model 1701 and convolved, and the extracted features are input to supervised classification model 1702. Supervised training of supervised classification model 1702 is then performed using backpropagation so that a loss function based on the error between the classification data output from supervised classification model 1702 and the training data is minimized.

[0084] Furthermore, when training distance learning model 1703, training of self-learning model 1701 is stopped, supervised data is input to self-learning model 1701 and convolution is performed, and the extracted features are input to distance learning model 1703. Then, distance learning of distance learning model 1703 is performed by the backpropagation method using basic loss functions such as contrastive loss and triplet loss.

[0085] D. Application Examples According to the information processing system 100 to which the present disclosure is applied, it is possible to support a user's learning by utilizing time-series media information that represents the actions and behaviors of a user who is studying. Examples of time-series media information include video, audio, and text, and information that can be recognized based on signals and sensor data that can be sensed by the sensor unit 101. For example, sensor data acquired by a swing sensor attached to a golf club or bat, and biosignals acquired by an IMU (Inertial Measurement Unit) or biosensor attached to a user during sports training can also be used as time-series media information.

[0086] FIG. 19 summarizes the relationship between the content of learning that can be supported by the present disclosure and the type of time-series media information.

[0087] When the present disclosure is applied to language learning, the information processing system 100 can be used to support the user's learning by using time-series media information such as voice signals spoken by the user, sentences obtained by voice recognition, and sentences written by the user.

[0088] When the present disclosure is applied to playing an instrument, the information processing system 100 can be used to assist the user in learning to play an instrument by using audio signals produced by the instrument being played by the user and video footage of the user playing as time-series media information.

[0089] When the present disclosure is applied to speeches or public speaking, the information processing system 100 can be used to support the improvement of a user's speech or public speaking skills by using time-series media information such as audio signals uttered by the user, text or manuscripts obtained by voice recognition of the user's speech, and video footage of the user giving a speech or public speaking.

[0090] When the present disclosure is applied to training for golf, baseball, or other sports, videos taken of a user during training can be used as time-series media information to support the user's training using the information processing system 100. Although omitted in Fig. 19, sensor data acquired by a swing sensor attached to a golf club or bat, and biosignals acquired by an IMU or biosensor attached to a user during sports training can also be used as time-series media information.

[0091] When the present disclosure is applied to cooking, a video of a user cooking can be used as time-series media information to assist the user in cooking using the information processing system 100.

[0092] When the present disclosure is applied to surgery or other medical procedures, as well as various treatments such as massage, videos of the user during surgery, examination, or treatment can be used as time-series media information to support the user in improving their medical or treatment skills using the information processing system 100.

[0093] When the present disclosure is applied to a user's writing activities such as writing novels, screenplays, and translations, the information processing system 100 can be used to support the user's writing skills by using the text written by the user as time-series media information.

[0094] When the present disclosure is applied to acting in movies or dramas, or stand-up comedy, the information processing system 100 can be used to support the user's acting by using video footage of the user acting, the user's spoken voice, voice-recognized sentences, or script sentences as time-series media information.

[0095] E.UI Example According to the information processing system 100 to which the present disclosure is applied, the determination unit 103 determines the learning level of a user's actions and actions based on time-series media information such as video and audio representing the user's actions and actions, and if it determines that there is an area of ​​insufficient learning, it can determine which areas of the user's actions and actions are insufficient and to what extent.The presentation unit 104 can visually present, as feedback to the user, areas in the time-series media information where the user's actions and actions differ from a reference action or action.Furthermore, by setting information about the reference action or action in advance, the user can conduct learning that is preferable to them.

[0096] This section describes an example of the configuration of a UI (User Interface) for presenting the user with the results of the assessment of the user's learning level. The screen for displaying the UI is assumed to be, for example, the screen of a personal computer or smartphone equipped with at least some of the components of the information processing system 100.

[0097] 20 shows an example of the configuration of a UI screen that presents the judgment result for a user's utterance when the user is learning a second language. This UI screen displays the speech waveform signal when the user utters the phrase "This was easy for us" along with the character string "This was easy for us."

[0098] When the determination unit 103 determines that the phrase "This was easy for us" spoken by the user differs from that of a native speaker and that the user's learning is insufficient, the determination unit 103 determines which parts of the speech waveform signal are insufficient and to what extent. Then, as shown in FIG. 20 , the presentation unit 104 highlights the parts of the speech waveform signal that differ from that of a native speaker, and also highlights the words or character strings "This," "eas," and "for" whose pronunciation is determined to differ from that of a native speaker. Furthermore, by displaying a moving image of the face of a native speaker when speaking "This was easy for us" on the UI screen, the user can learn mouth movements to achieve pronunciation similar to that of a native speaker. Of course, a UI with a screen layout other than that shown in FIG. 20 may be used as long as it can present information such as the degree to which the user's speech waveform signal differs from that of a native speaker and the words or character strings in the phrase spoken by the user that differ from that of a native speaker. Furthermore, for example, the user can set their ideal pronunciation by previously setting the place of origin of the reference native speaker (e.g., whether British English or American English), age, social class, etc.

[0099] 21 shows an example of the configuration of a UI screen that presents the evaluation results of a user's violin performance while the user is practicing playing the violin. This UI screen displays the audio waveform signals emitted from the violin played by the user along with the musical score that was played.

[0100] If the determination unit 103 determines that the user's violin performance differs from that of a professional violinist and that the user has insufficient learning in some areas, it determines which areas and to what extent the user's learning is insufficient on the audio waveform signal. Then, as shown in FIG. 21 , the presentation unit 104 highlights the areas on the audio waveform signal that differ from that of the professional violinist, and also highlights the notes on the score that differ from that of the professional violinist. Of course, a UI with a screen layout other than that shown in FIG. 21 may be used as long as it can present information such as the extent to which the audio waveform signal of the user's violin performance differs from that of a professional violinist, and the notes on the score that differ from that of the professional violinist. Furthermore, for example, the user can set their ideal performance by previously setting the instrument, style, or playing style of a professional violinist to be used as a reference.

[0101] 22 shows an example of the configuration of a UI screen that displays the results of a user's bat swing while the user is practicing. This UI screen displays multiple still images, framed out at predetermined time intervals from a video of the user swinging the bat, arranged in chronological order, along with waveform signals that represent the amount of displacement for each of the user's major body parts, such as the forearm, hand, knee, and toe.

[0102] If the determination unit 103 determines that the user's bat swing differs from that of a professional baseball player and that the user has not practiced enough, the determination unit 103 determines which parts of the user's body differ from that of the professional baseball player and to what extent. Then, as shown in FIG. 22 , the presentation unit 104 highlights body parts that differ from that of the professional baseball player on each still image arranged in chronological order, and also highlights sections of the displacement signal of each body part that differ from that of the professional baseball player. Of course, a UI with a screen layout other than that shown in FIG. 22 may be used as long as it can visually represent how and to what extent the user's bat swing differs from that of the professional baseball player. Furthermore, for example, the user can set their ideal swing by previously setting the age, physique, playing style, etc. of a reference professional baseball player.

[0103] F. Equipment configuration example F-1. Example of information processing device configuration Fig. 23 shows an example of the configuration of an information processing device 2300. The information processing device 2300 corresponds to, for example, a general personal computer, but can operate as the device 700 shown in Fig. 7, the second device 802 shown in Fig. 8, the second device 902 shown in Fig. 9, and the second device 1002 shown in Fig. 10. Each element of the information processing device 2300 will be described below.

[0104] A CPU (Central Processing Unit) 2301 is interconnected with a ROM (Read Only Memory) 2302, a RAM (Random Access Memory) 2303, a hard disk drive (HDD) 2304, and an input / output interface 2305 via a bus 2310.

[0105] The CPU 2301 executes programs loaded from the ROM 2302 or HDD 2304 to the RAM 2303, and can perform various processes while temporarily storing working data in the RAM 2303 during execution. The programs executed by the CPU 2301 include a basic input / output program stored in the ROM 2302, and an operating system (OS) and application programs installed in the HDD 2304. The OS provides an execution environment for the application programs. The application programs include a learning support application program that determines the user's learning level based on sensor information (time-series media information).

[0106] The ROM 2302 is a read-only memory that permanently stores basic input / output programs, device information, etc. The RAM 2303 is composed of volatile memory such as DRAM (Dynamic RAM), and is used as a working area for the CPU 2301. The HDD 2304 is a large-capacity storage device that uses one or more magnetic disks fixed within the unit as recording media, and stores programs and data in file format. An SSD (Solid State Drive) may be used instead of the HDD.

[0107] The input / output interface 2305 is connected to various input / output devices, such as an output unit 2311, an input unit 2312, a communication unit 2313, and a drive 2314. The output unit 2311 is composed of a display device such as an LCD (Liquid Crystal Display), a speaker, a printer, or other output device, and outputs, for example, the results of program execution by the CPU 1601. The display device can be used to present the results of determining the user's learning level. The input unit 2312 is composed of a keyboard, mouse, touch panel, or other input device, and accepts instructions from the user. The input unit 2312 also includes a microphone, camera, and other sensors, and acquires time-series media information, such as video and audio, related to the user's movements and actions. The output unit 2311 and the input unit 2312 may also have interfaces, such as USB or HDMI (registered trademark), for externally connecting external output and input devices.

[0108] The communication unit 2313 has a wired or wireless communication interface conforming to a predetermined communication protocol, and performs data communication with an external device. An example of a wired communication interface is Ethernet (registered trademark). An example of a wireless communication interface is Wi-Fi (registered trademark) or Bluetooth (registered trademark). When the information processing device 2300 operates as a second device, the communication unit 2313 communicates with a first device.

[0109] The communication unit 2313 is also connected to a wide area network such as the Internet. Using the communication unit 2313, it is possible to download, for example, an application program (described above) from a download site on the Internet and install it in the information processing device 2300.

[0110] The drive 2314 has the removable recording medium 1615 loaded therein and performs read processing from the removable recording medium 2315 and write processing to the removable recording medium 2315 (if the removable recording medium is writable). The removable recording medium 2315 stores programs, data, and the like in file format. For example, a removable recording medium 2315 storing packaged software such as an application program (described above) can be loaded into the drive 2314 and installed in the computer 2300. Examples of the removable recording medium 2315 include a flexible disk, a CD-ROM (Compact Disc Read Only Memory), an MO (Magneto Optical) disk, a DVD (Digital Versatile Disc), a magnetic disk, and a semiconductor memory.

[0111] F-2. Example of information terminal configuration Fig. 24 shows a configuration example of an information terminal 2400. The information terminal 2400 corresponds to a multi-function information terminal such as a smartphone or a tablet, and can operate as the device 700 shown in Fig. 7, the first device 801 shown in Fig. 8, the first device 901 shown in Fig. 9, and the first device 1001 shown in Fig. 10.

[0112] The information terminal 2400 includes a built-in antenna 2401, a mobile communication processing unit 2402, a microphone 2403, a speaker 2404, a memory unit 2405, an operation unit 2406, a display unit 2407, a control unit 2408, a control line 2409, a data line 2410, a WLAN communication antenna 2411, a WLAN communication control unit 2412, a BLE (Bluetooth (registered trademark) Low Energy) communication antenna 2413, a BLE communication control unit 2414, an infrared transmission / reception unit 2415, a non-contact communication antenna 2416, a non-contact communication control unit 2417, a GNSS (Global Navigation Satellite System) receiving antenna 2418, a GNSS positioning unit 2419, a camera unit 2420, a memory slot 2421, and a sensor unit 2423. Each of the components of the information terminal 2400 will be described below.

[0113] The built-in antenna 2401 is configured to receive signals transmitted through a mobile phone network such as LTE or NR, and to transmit signals from the device itself to the mobile phone network. The mobile communication processing unit 2402 demodulates and decodes signals received by the built-in antenna 2401, and encodes and modulates transmission data to be transmitted to the mobile phone network via the built-in antenna 2401.

[0114] The microphone 2403 collects audio, converts it into an electrical signal, and then performs AD conversion. The audio signal digitized by the microphone 103 is supplied to the mobile communication processing unit 2402 via the data line 2410, where it is encoded and modulated before being sent to the mobile phone network via the built-in antenna 2401. The microphone 2403 mainly functions as a transmitter, but in this embodiment, it also functions as the sensor unit 101 that collects the user's speech and acquires an audio waveform signal (time-series media information).

[0115] The speaker 2404 mainly functions as a receiver, and converts the digital audio signal supplied from the mobile communication processing unit 2402 via the data line 2410 into digital audio, and then performs amplification processing and other processes before emitting the audio.

[0116] The memory unit 2405 includes a ROM, a RAM, and a nonvolatile memory such as an EEPROM (Electrically Erasable and Programmable ROM) or a flash memory.

[0117] The ROM stores various programs (applications) such as various program codes executed by a CPU (Central Processing Unit) constituting the control unit 2408 (described later), program code for operating an email process for editing emails, and a program for processing images captured by the camera unit 120, as well as important data such as identification information (identification ID) of the mobile phone terminal and data required for various processes. The RAM is mainly used as a working area, for example, to temporarily store intermediate results while the CPU is executing various processes.

[0118] The nonvolatile memory stores and holds data that should be retained even when the power to the information terminal 2400 is turned off. Examples of data stored and held in the nonvolatile memory include address book data, email data, image data of images captured by the camera unit 2420, various web data such as image data and text data downloaded via the Internet, various setting parameters, dictionary information, and additional programs.

[0119] The operation unit 2406 is composed of, for example, a touch panel superimposed on the screen of the display unit 2407 (described later), numeric keys, several symbol keys, several function keys, a so-called jog dial key that can be rotated and pressed, etc. The operation unit 2406 accepts operation input from the user of the information terminal 2400, converts this into an electrical signal, and supplies it to the control unit 2408 via a control line 2409. This enables the control unit 2408 to control each unit in accordance with instructions from the user and perform processing in accordance with the user's instructions.

[0120] The display unit 2407 is composed of a thin display element such as an organic EL (Electro Luminescence) or LCD and its control circuit, and displays various information supplied via a control line 2409. For example, it can display information such as various image data and e-mail data captured via the built-in antenna 2401 and mobile communication processing unit 2402, text data input via the operation unit 2406, prepared operation guidance and various message information, and image data captured via the camera unit 2420. If the operation unit 2406 includes a touch panel superimposed on the screen of the display unit 2407, the user can directly perform input operations on display objects on the screen.

[0121] The control unit 2408 is a main controller that performs overall control of the information terminal 2400. Specifically, the control unit 2408 is configured with a CPU, and loads programs stored in the ROM of the memory unit 2405 into the RAM, executes them, generates control signals to be supplied to each unit, and passes them to each unit via a control line 2409. The programs executed by the control unit 108 include a program (application) that performs processing related to determining the user's learning level. Furthermore, the control unit 2408 receives information provided from each unit, generates new control signals in response to the information, and supplies them via the control line 2409.

[0122] The control line 2409 is a bus line mainly for transmitting control signals and various information associated with control, while the data line 2410 is a bus line for transmitting various data to be sent and received, such as audio data, image data, and e-mail data, as well as various data to be processed.

[0123] The WLAN communication antenna 2411 is configured to receive signals transmitted through a WLAN that uses an unlicensed band, such as the 2.4 GHz band or the 5 GHz band, and to transmit signals from the device itself to the WLAN. The WLAN communication control unit 2412 controls WLAN communication operations that use the unlicensed band, and also performs demodulation and decoding of signals received by the WLAN communication antenna 2411, and encoding and modulation of transmission data that is sent to the WLAN via the WLAN communication antenna 2411. The WLAN communication control unit 112 controls one-to-one wireless communication in ad hoc mode, and wireless communication in infrastructure mode that connects to a nearby access point and then to a WLAN.

[0124] The BLE communication antenna 2413 is configured to transmit and receive BLE signals. The BLE communication control unit 2414 controls the BLE communication operation, and also performs demodulation and decoding processes on signals received by the BLE communication antenna 2411, and encoding and modulation processes on transmission data sent via the BLE communication antenna 2411.

[0125] The infrared transmitter / receiver 2415 includes an LED (Light Emitting Diode) for emitting infrared light and a photodetector for receiving infrared light, and transmits and receives signals using infrared light in a frequency band slightly lower than that of visible light. By bringing the information terminal 2400 close to another terminal through this infrared transmitter / receiver 2415, data such as e-mail addresses and images can be exchanged by transmitting and receiving infrared light. Since infrared communication is performed between mobile phone terminals in close proximity, communication can be performed while maintaining security.

[0126] The contactless communication antenna 2416 is configured to transmit, receive, or transmit and receive contactless signals using electromagnetic induction. The contactless communication control unit 2417 controls contactless communication operations using contactless communication technology such as FeliCa (registered trademark). Specifically, the contactless communication control unit 2417 controls operations as a card, reader, or reader / writer in a contactless communication system.

[0127] The GNSS receiving antenna 2418 and the GNSS positioning unit 2419 analyze the GNSS signals received from GNSS satellites to identify the current position of the information terminal 2400. Specifically, the GNSS receiving antenna 2418 receives GNSS signals from multiple GNSS satellites, and the GNSS positioning unit 2419 synchronizes, demodulates, and analyzes each GNSS signal received through the GNSS receiving antenna 2418 to calculate position information. The information on the current position calculated by the GNSS positioning unit 2419 is used, for example, for navigation functions and for metadata indicating the shooting position that is added to image data captured by a camera unit 2420 (described later).

[0128] Although not shown, the information terminal 2400 further includes a clock circuit that provides the current date, current day of the week, and current time. The current date and time acquired from this clock circuit is added to image data captured by the camera unit 2420 (described later) as metadata indicating the date and time of capture.

[0129] The camera unit 2420 includes an objective lens, a shutter mechanism, and an imaging element such as a CMOS (Complementary Metal Oxide Semiconductor) (none of which are shown). The imaging element captures an electrical signal of the image of the subject, converts it into digital data, and supplies it to the memory unit 2405 via a data line 2410 for recording. In this embodiment, the camera unit 2420 also functions as a sensor unit 101 that captures the movements and actions of the user during learning and acquires video (time-series media information).

[0130] The memory slot 2421 is a device into which an external memory 2422 such as a removable microSD card is inserted. For example, the user can use the external memory 2422 as a user memory when the available storage capacity in the memory unit 2405 becomes insufficient, or can insert the external memory 2422, in which a program (application) for realizing a new function is recorded, into the memory slot 2421 to add the new function to the information terminal 2400.

[0131] The sensor unit 2423 may include other sensor elements such as an illuminance sensor, an IMU (Inertial Measurement Unit), a TOF (Time Of Flight) sensor, a temperature sensor, a humidity sensor, etc. Note that the microphone 2403 can be regarded as a sound sensor, the GNSS communication control unit 2419 as a positioning sensor, and the camera unit 2420 as an image sensor, and these can be treated as parts of the sensor unit 2423. [Industrial Applicability]

[0132] Although the present disclosure has been described in detail above with reference to specific embodiments, it is obvious that those skilled in the art can make modifications or substitutions to the embodiments without departing from the spirit and scope of the present disclosure.

[0133] Although the present specification has mainly described embodiments in which the present disclosure is applied to speech learning using user voice input, the gist of the present disclosure is not limited thereto. For example, the present disclosure can also be applied to learning physical movements using video of a user captured by a camera, or learning using a combination of audio and video. Furthermore, the present disclosure can be applied not only to speech in a second language but also to learning sentences.

[0134] Furthermore, the content described herein as a language learning process based on native speaker pronunciation may also be implemented as a process for reducing the deviation between the user's pronunciation and the reference pronunciation, for example, a process based on the difference between the standard accent of a language and the regional accent of the user.

[0135] Furthermore, the application field of the present disclosure is not limited to language learning. For example, the present disclosure can be similarly applied to learning to play a musical instrument, learning speeches and public speaking (studying audio, video (body movements), and writing), learning various sports using videos, learning cooking using videos, learning various treatments such as surgery and other medical procedures and massage using videos, learning writing in writing activities (novels, scripts, translations, etc.), and learning acting (speech, writing, body movements, etc.) by actors and comedians.

[0136] In short, the present disclosure has been described in the form of examples, and the contents of the specification should not be interpreted as limiting. To determine the gist of the present disclosure, the claims should be taken into consideration.

[0137] The present disclosure may also be configured as follows.

[0138] (1) an input step of inputting time-series media information representing a user's actions or behavior during learning; a first determination step of determining a learning level of a user based on time-series media information; an output step of outputting a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior based on the learning level of the user determined in the first determination step; An information processing method comprising:

[0139] (2) in the output step, a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior is output to a presentation device; The method further includes a first presentation step of presenting to the user, via the presentation device, a portion of the time-series media information where the user's action or behavior differs from a reference action or behavior. The information processing method according to (1) above.

[0140] (3) in the first presentation step, a portion where the user's action or behavior differs from a reference action or behavior is visually presented on the time-series information medium; The information processing method described in (2) above.

[0141] (4) In the first presentation step, the user is visually presented with a word or character in the phrase spoken by the user that differs from the standard (or ideal) pronunciation of a native speaker. The information processing method described in (3) above.

[0142] (5) In the first presentation step, a part of the user's body motion that differs from a reference body motion (or an ideal body motion of a professional athlete or trainer) is visually presented. The information processing method described in (3) above.

[0143] (6) a second determination step of determining distance information representing a difference between the user's movement or action and a reference movement or action; The method further includes a second presentation step of outputting the determination result of the second determination step to a presentation device and presenting it to a user. The information processing method according to any one of (1) to (5) above.

[0144] (7) In the second presentation step, the distance information is visually presented in an N-dimensional space with the reference action or behavior positioned at the center. The information processing method according to (6) above.

[0145] (8) continuously executing the input step and the first determination step until it is determined that the user's learning level is sufficient; The information processing method according to any one of (1) to (7) above.

[0146] (9) In the first determination step, a determination is made using a trained machine learning model. The information processing method according to any one of (1) to (8) above.

[0147] (10) The first determination step processes the time-series media information using a self-trained first machine learning model, and then determines the user's learning level using a supervised second machine learning model. The information processing method according to (9) above.

[0148] (11) In the first determination step, an attention mechanism incorporated in the second machine learning model is used to determine, based on time-series media information, the basis for determining that the user's learning level is insufficient or that the user needs to learn. The information processing method according to (10) above.

[0149] (12) The machine learning model includes a feature extraction unit that extracts features of the time-series media information and a classification unit that classifies learning levels based on the extracted features. The information processing method according to (9) above.

[0150] (13) The feature extraction unit is trained by self-learning, and the classification unit is trained by supervised learning using the trained feature extraction unit. The information processing method according to (12) above.

[0151] (14) In the first determination step, an attention mechanism incorporated in the classification unit is used to determine, based on the time-series media information, the basis for determining that the user's learning level is insufficient or that the user needs to learn. The information processing method according to any one of (12) and (13) above.

[0152] (15) A second determination step of determining distance information representing a difference between a user's movement or action and a reference movement or action from the feature amount of the time-series media information extracted by the feature extraction unit is further included. The information processing method according to any one of (12) to (14) above.

[0153] (16) In the second determination step, distance information is determined by a distance learning model using contrastive loss or triplet loss as a loss function. The information processing method according to (15) above.

[0154] (17) an input unit for inputting time-series media information representing the user's actions or behavior during learning; a first determination unit that determines a learning level of a user based on time-series media information; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior based on the learning level of the user determined by the first determination unit; An information processing system comprising:

[0155] (18) The present invention further includes a sensor unit that detects the user's movements or actions during learning and acquires time-series media information, and a presentation device that outputs the parts of the time-series media information where the user's movements or actions differ from a reference movement or action. The information processing system according to (17) above.

[0156] (19) A sensor unit that detects the user's actions or behavior during learning and acquires time-series media information; a communication unit that transmits the time-series media information to an external device and receives from the external device a determination result of the user's learning level and a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior; a presentation unit that presents the received information; An information terminal equipped with the above.

[0157] (20) an input unit for inputting time-series media information representing the user's actions or behavior during learning; a first determination unit that determines a learning level of a user based on time-series media information; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior, based on the learning level of the user determined by the first determination unit; A computer program written in a computer-readable format to cause a computer to function as a [Explanation of symbols]

[0158] 100...information processing system, 101...sensor unit, 102...input unit 103...judgment section, 104...presentation section 1400...DNN, 1410...Feature extraction unit, 1411...CNN 1420...Classification section, 1421...Bi-LSTM layer 1422...Attention mechanism, 1423...Classification unit 1701...Self-learning model, 1702...Supervised classification model 1703…Distance Learning Model 1801...Encoder section, 1802...Transformer section 2300...information processing device, 2301...CPU, 2302...ROM 2303...RAM, 2304...HDD 2305...input / output interface, 2310...bus 2311...output unit, 2312...input unit, 2313...communication unit 2314...Drive, 2315...Removable recording medium 2400...information terminal, 2401...built-in antenna 2402...mobile communication processing unit, 2403...microphone, 2404...speaker 2405...Memory section, 2406...Operation section, 2407...Display section 2408...control section, 2409...control line, 2410...data line 2411...WLAN communication antenna, 2412...WLAN communication control unit 2413...BLE communication antenna, 2414...BLE communication control unit 2415...infrared transmitting / receiving unit, 2416...non-contact communication antenna 2417...Non-contact communication control unit, 2418...GNSS receiving antenna 2419...GNSS positioning unit, 2420...camera unit 2421...Memory slot, 2422...External memory 2423...Sensor section

Claims

1. an input step of inputting time-series media information representing the user's actions or behavior during learning; a first determination step of processing time-series media information using a first machine learning model that has been self-trained using a large amount of data collected through broadcasting or distribution services, etc., and then determining the user's learning level using a second machine learning model that has been supervised-trained using the user's behavior data and ideal behavior data; an output step of outputting a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior, based on the learning level of the user determined in the first determination step; An information processing method comprising:

2. an input step of inputting time-series media information representing the user's actions or behavior during learning; a first determination step of determining a learning level of a user using a trained machine learning model including a feature extraction unit that has been self-trained to extract features of time-series media information and a classification unit that has been trained by supervised learning using the feature extraction unit that has been trained to classify learning levels based on the extracted features; an output step of outputting a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior, based on the learning level of the user determined in the first determination step; An information processing method comprising:

3. an input step of inputting time-series media information representing the user's actions or behavior during learning; a first determination step of determining a learning level of a user using a trained machine learning model including a feature extraction unit that extracts features of time-series media information and a classification unit that classifies learning levels based on the extracted features; an output step of outputting a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior, based on the learning level of the user determined in the first determination step; a second determination step of determining distance information representing a difference between a user's movement or action and a reference movement or action from the feature amount of the time-series media information extracted by the feature extraction unit using a trained distance learning model that has been trained to determine distance information between the user's movement or action and the reference movement or action from the time-series media information obtained by sensing the user's movement or action; An information processing method comprising:

4. In the output step, a portion of the time-series media information in which the user's action or behavior differs from a reference action or behavior is output to a presentation device; a first presentation step of presenting to the user, via the presentation device, a portion of the time-series media information where the user's action or behavior differs from a reference action or behavior; 4. The information processing method according to claim 1.

5. In the first presenting step, a portion of the time-series media information where the user's action or behavior differs from a reference action or behavior is visually presented. The information processing method according to claim 4.

6. In the first presentation step, a word or character in a phrase spoken by the user that differs from a standard (or ideal) pronunciation of a native speaker is visually presented. The information processing method according to claim 5 .

7. In the first presentation step, a part of the user's body motion that differs from a reference body motion (or an ideal body motion of a professional athlete or trainer) is visually presented. The information processing method according to claim 5 .

8. The method further includes a second presentation step of outputting the determination result of the second determination step to a presentation device and presenting the result to a user. The information processing method according to claim 3 .

9. In the second presentation step, the distance information is visually presented in an N-dimensional space in which a reference action or behavior is placed at the center. The information processing method according to claim 8.

10. The input step and the first determination step are continuously performed until it is determined that the user's learning level is sufficient.

4. The information processing method according to claim 1.

11. In the first determination step, an attention mechanism incorporated in the second machine learning model is used to determine, based on time-series media information, a basis for determining that the user's learning level is insufficient or that the user needs to learn. The information processing method according to claim 1 .

12. In the first determination step, an attention mechanism incorporated in the classification unit is used to determine, based on the time-series media information, a basis for determining that the user's learning level is insufficient or that the user needs to learn. The information processing method according to claim 2 .

13. In the second determination step, distance information is determined by a distance learning model using Contrastive Loss or Triplet Loss as a loss function. The information processing method according to claim 3 .

14. an input unit for inputting time-series media information representing the actions or behavior of a user during learning; a first determination unit that processes time-series media information using a first machine learning model that has been self-trained using a large amount of data collected through broadcasting or distribution services, and then determines the learning level of the user using a second machine learning model that has been supervised and trained using the user's motion data and ideal motion data; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior based on the learning level of the user determined by the first determination unit; An information processing system comprising:

15. an input unit for inputting time-series media information representing the actions or behavior of a user during learning; a first determination unit that determines a learning level of a user using a trained machine learning model including a feature extraction unit that has been self-trained to extract features of time-series media information, and a classification unit that has been trained by supervised learning using the feature extraction unit that has been trained to classify learning levels based on the extracted features; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior based on the learning level of the user determined by the first determination unit; An information processing system comprising:

16. an input unit for inputting time-series media information representing the actions or behavior of a user during learning; a first determination unit that determines a learning level of a user using a trained machine learning model including a feature extraction unit that extracts features of time-series media information and a classification unit that classifies learning levels based on the extracted features; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior based on the learning level of the user determined by the first determination unit; a second determination unit that determines distance information representing a difference between a user's movement or action and a reference movement or action from the feature amount of the time-series media information extracted by the feature extraction unit, using a trained distance learning model that has been trained to determine distance information between the user's movement or action and the reference movement or action from the time-series media information obtained by sensing the user's movement or action; An information processing system comprising:

17. The present invention further includes a sensor unit that detects the user's actions or behavior during learning and acquires time-series media information, and a presentation device in which the output unit outputs a portion of the time-series media information in which the user's actions or behavior differ from a reference action or behavior.

17. An information processing system according to any one of claims 14 to 16.

18. an input unit for inputting time-series media information representing the actions or behavior of a user during learning; a first determination unit that processes time-series media information using a first machine learning model that has been self-trained using a large amount of data collected through broadcasting or distribution services, and then determines the user's learning level using a second machine learning model that has been supervised and trained using the user's behavior data and ideal behavior data; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior, based on the learning level of the user determined by the first determination unit; A computer program written in a computer-readable format to cause a computer to function as a

19. an input unit for inputting time-series media information representing the actions or behavior of a user during learning; a first determination unit that determines the learning level of a user using a trained machine learning model including a feature extraction unit that has been self-trained to extract feature quantities of time-series media information and a classification unit that has been trained by supervised learning using the feature extraction unit that has been trained to classify learning levels based on the extracted feature quantities; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior, based on the learning level of the user determined by the first determination unit; A computer program written in a computer-readable format to cause a computer to function as a

20. an input unit for inputting time-series media information representing the actions or behavior of a user during learning; a first determination unit that determines a learning level of a user using a trained machine learning model including a feature extraction unit that extracts feature amounts of time-series media information and a classification unit that classifies learning levels based on the extracted feature amounts; an output unit that outputs a portion of the time-series media information in which a user's action or behavior differs from a reference action or behavior, based on the learning level of the user determined by the first determination unit; a second determination unit that determines distance information representing a difference between a user's movement or action and a reference movement or action from the feature amount of the time-series media information extracted by the feature extraction unit, using a trained distance learning model that has been trained to determine distance information between the user's movement or action and the reference movement or action from the time-series media information obtained by sensing the user's movement or action; A computer program written in a computer-readable format to cause a computer to function as a

Citation Information

Patent Citations

  • Utterance evaluation device, utterance evaluation method, and program

    JP2016090900A

  • Information processing apparatus, data classification method and program

    JP2020042330A

  • Data processing device, data processing method and data processing program

    JP2020149601A

  • Learning device, evaluation device and production method for leaning model

    JP2021001959A

  • Voice learning system and method for voice learning

    JP2021113904A