Online lesson coaching support device

The online lesson coach support device addresses the challenge of remote coaching by analyzing user emotions and exercise form to generate personalized instruction scenarios, improving the quality of online coaching.

JP7840230B2Active Publication Date: 2026-04-03NTT DOCOMO INC
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-08-10
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In online lessons, coaches face challenges in accurately judging the user's situation and providing appropriate guidance due to the limitations of remote communication.

Method used

An online lesson coach support device that analyzes user emotions and exercise form through video and audio signals, generating personalized instruction scenarios based on emotion estimation and exercise form evaluation to support coaches in providing effective guidance.

Benefits of technology

Enables coaches to provide timely and appropriate instructions to users during online lessons, enhancing the effectiveness of remote coaching by leveraging real-time emotion and form analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007840230000001
    Figure 0007840230000001
  • Figure 0007840230000002
    Figure 0007840230000002
  • Figure 0007840230000003
    Figure 0007840230000003
Patent Text Reader

Abstract

To support a coach to properly teach a user in an online lesson.SOLUTION: A coach support device for an on-line lesson which receives image signals of a user and voice signals of the user during the lesson from a terminal device via a network comprises: an estimation unit 112 which estimates a user's emotion based on the image signals and voice signals; an evaluation unit 113 which evaluates an exercise form of the user based on the image signals; and a teaching scenario generation unit 114 which generates a teaching scenario for the user based on at least a result of the emotion evaluation by the estimation unit 112 and a result of the exercise form evaluation by the evaluation unit 113.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to, for example, a coach support device for online lessons.

Background Art

[0002] For health, yoga, dance, injury prevention, etc., it is common for users to go to a gym or studio and actually have lessons facing a coach. However, it is a heavy burden for users to go to a gym or studio. Therefore, in recent years, online lessons have been proposed in which users have lessons with a coach online, for example, via video call (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in online lessons, situations may occur where it is difficult for a coach to judge the user's situation, or situations may occur where the coach cannot convey appropriate guidance content to the user. In view of such circumstances, an object of the present invention is to support a coach so that the coach can appropriately guide a user in an online lesson.

Means for Solving the Problems

[0005] To solve the above problems, an online lesson coach support device according to one aspect of the present disclosure is an online lesson coach support device that receives a video signal and an audio signal of a user during a lesson from a terminal device via a network, and comprises: an estimation unit that estimates the user's emotions based on the video signal and the audio signal; an evaluation unit that evaluates the user's exercise form based on the video signal; and an instruction scenario generation unit that generates an instruction scenario for the user based on at least the emotion estimation result by the estimation unit and the exercise form evaluation result by the evaluation unit. According to the coach support device for online lessons in the above-described embodiment, it is possible to support coaches so that they can appropriately instruct users in online lessons. [Brief explanation of the drawing]

[0006] [Figure 1] This diagram shows the schematic configuration of the online lesson system. [Figure 2] This diagram shows the configuration of the coach support device in the system. [Figure 3] This block shows the functions built into the coach support device. [Figure 4] This diagram shows the configuration of terminal devices in the system. [Figure 5] This diagram shows the processing and data flow in the estimation and evaluation units. [Figure 6] This diagram shows the processing and data flow in the instruction scenario generation unit. [Figure 7] This is a flowchart showing the operation of the coaching support device. [Figure 8] This figure shows an example of a screen displayed by a coaching support device. [Figure 9] This figure shows an example of a screen displayed on a terminal device. [Figure 10] This figure shows another example of a screen displayed on a coaching support device. [Figure 11] This figure shows an example of a screen displayed on a terminal device. [Modes for carrying out the invention]

[0007] Hereinafter, embodiments for carrying out the present invention will be described with reference to the drawings. Figure 1 is a schematic diagram of an online lesson system 1 including a coach support device 10 according to an embodiment. In the online lesson system 1, the coach support device 10 and a plurality of terminal devices 20-1, 20-2, ..., 20-n are connected via a network 30. In this embodiment, online lessons are possible not only between a user and a coach one-on-one, but also between multiple users and coaches (n-to-n).

[0008] The coach support device 10 is, for example, a personal computer and is operated by a coach who instructs the user in an online lesson. The coach support device 10 has a program installed for providing online lessons.

[0009] Terminal devices 20-1, 20-2, 20-3, ..., 20-n are, for example, smartphones or personal computers, and are operated by users taking online lessons. In the following, to describe a typical terminal device, terminal devices 20-1, 20-2, 20-3, ..., 20-n will be simply referred to as "20," omitting the hyphen for identification. Each terminal device 20 has a program (browser) installed for taking online lessons.

[0010] Figure 2 shows the hardware configuration of the coach support device 10. The coach support device 10 includes a processing device 110, a storage device 120, a display device 130, a communication device 140, a camera 150, and an audio input / output device 160. Although not specifically shown in the diagram, the coach support device 10 has an operation input device such as a keyboard that accepts operations entered by the coach.

[0011] The processing device 110 is composed of, for example, one or more processors and controls each element of the coach support device 10. Specifically, the processing device 110 is composed of one or more types of processors such as a CPU (Central Processing Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit).

[0012] The storage device 120 stores the programs executed by the processing device 110 and various data used by the processing device 110. As the storage device 120, for example, known recording media such as semiconductor recording media and magnetic recording media, or a combination of multiple types of recording media is adopted. The display device 130 is composed of, for example, a liquid crystal display panel or an organic EL panel, and displays videos including various images under the control of the processing device 110. Specifically, in an online lesson, the display device 130 displays the videos of one or more users, the video of the coach, and the advice to the users.

[0013] The communication device 140 communicates with each of the terminal devices 20 under the control of the processing device 110. The imaging device 150 is a camera that captures the entire body of the coach in an online lesson. The audio input / output device 160 is composed of a microphone that inputs the voice of the coach, a speaker that outputs the voice from the user, and the like.

[0014] FIG. 3 is a block diagram showing a functional configuration constructed by the processing device 110 and the storage device 120 in the coach support device 10. The processing device 110 constructs a plurality of elements (control management unit 111, estimation unit 112, evaluation unit 113, guidance scenario generation unit 114, and presentation unit 115) by executing the programs stored in the storage device 120.

[0015] The control and management unit 111 manages the control and progress of each part in the online lesson. The estimation unit 112 analyzes the video, voice, etc. of the user transmitted from the terminal device, and estimates the emotion of the user. The evaluation unit 113 analyzes the video signal transmitted from the terminal device, estimates the posture of the user, and evaluates the movement form of the user based on the estimated posture of the user.

[0016] The guidance scenario generation unit 114 generates a guidance scenario including advice to the user based on the user emotion estimation result by the estimation unit 112, the user movement form evaluation result by the evaluation unit 113, and further movement guidance data, movement menu data, guidance judgment rule data, user basic information, dialogue data, and coach information. Note that various data, information, etc. required for generating the guidance scenario will be described later. Also, the guidance scenario includes not only advice to the user, but also information for supporting the coach in the online lesson, such as information indicating what the coach should do in the online lesson.

[0017] The presentation unit 115 controls the display content on the display device 230 and presents various contents to the coach. Specifically, the presentation unit 115 presents advice based on the guidance scenario to the coach, or presents which user the guidance scenario was generated for when there are multiple users as will be described later. Also, the presentation unit 115 causes the communication device 140 to transmit the voice signal and video signal of the coach to the terminal device 20. Thereby, on the terminal device 20, the voice and video of the coach are displayed on the display device 230.

[0018] The control and management unit 111 manages the movement guidance content, user basic information, and coach information using the database DB created in the storage device 120. The movement guidance content includes movement form data, movement guidance data, movement menu data, and guidance judgment rule data. Also, the database DB includes a learning database for analyzing the speech content of the user.

[0019] Figure 4 shows the hardware configuration of terminal device 12. Terminal device 20 includes a processing unit 210, a storage device 220, a display device 230, a communication device 240, a camera 250, and an audio input / output device 260.

[0020] The processing unit 210 controls each element of the terminal device 20. The storage device 220 stores programs executed by the processing unit 210 and various data used by the processing unit 210. The programs executed by the processing unit 210 include the operating system as well as programs for receiving online lessons. The display device 230 displays various types of images under the control of the processing device 210. Specifically, in online lessons, the display device 230 displays images from video calls with the coach, etc.

[0021] The communication device 240 communicates with the coach support device 10 under the control of the processing unit 210. The camera 250 is a camera that captures the user's entire body during online lessons. The audio input / output device 260 consists of a microphone for inputting the user's voice and a speaker for outputting the coach's voice. Although not specifically shown in the diagram, the terminal device 20 is equipped with a touch panel that is superimposed on the image display surface of the display device 230 and accepts user operations. The functional configuration of the processing unit 210 and the storage device 220 in the terminal device 20 is not important in this case, so its explanation will be omitted.

[0022] As described above, in this embodiment, it is possible to provide online lessons to multiple users simultaneously. However, for the sake of simplicity, Figures 5 to 8 illustrate the case where there is only one user of the online lesson.

[0023] Figure 5 is a diagram showing the flow of various data and various processes in the estimation unit 112 and the evaluation unit 113. In online lessons, users perform various exercise forms (poses / movements) while video conferring with their coach. Specifically, the voice spoken by the user is input to the audio input / output device 260 of the terminal device 20, and the user's entire body is captured by the camera device 250. The audio signal output from the audio input / output device and the video signal output from the camera 250 are transmitted to the coach support device 10 via the communication device 140 and the network 30.

[0024] In the coach support device 10, the received audio and video signals are transferred as follows: The audio signal is transferred to the estimation unit 112 as the user's voice Vo, and the video signal is transferred to the estimation unit 112 and the evaluation unit 113 as the user's video Vi.

[0025] The estimation unit 112 estimates the user's emotions based on the results of analyzing the user's voice Vo in four different ways and the results of analyzing the user's video Vi. The four types of analysis processes are language analysis process 112a, tone analysis process 112b, speed analysis process 112c, and timing analysis process 112d. The language analysis process 112a analyzes the user's voice Vo and identifies the meaning of the voice. The tone analysis process 112b analyzes the tone (pitch) of the user's voice Vo. For example, the tone analysis process 112b determines whether the voice tone is high / normal / low. The speed analysis process 112c analyzes the speed of the user's voice Vo. For example, the speed analysis process 112c determines whether the voice speed is fast / normal / slow.

[0026] Even if the semantic content identified by language analysis process 112a is the same, the actual meaning may differ depending on the tone and speed of the user's voice. For this reason, the semantic content identified by language analysis process 112a may be modified by the results of tone analysis process 112b and speed analysis process 112c.

[0027] The interval analysis process 112d analyzes the user's voice Vo to determine the continuity of the speech made by that voice. More specifically, the interval analysis process 112d determines whether, if the user's voice is temporally discrete, the discrete voices constitute a continuous speech.

[0028] The analysis process of the user's video Vi in the estimation unit 112 is the facial expression analysis process 112e. The facial expression analysis process 112e focuses on the user's face in the user's video Vi and analyzes the user's facial expressions.

[0029] The emotion estimation process 112f in the estimation unit 112 estimates the user's emotion based on at least one of the user's facial expression analyzed by the facial expression analysis process 112e and the user's statements indicated by the user's voice Vo. Specifically, the emotion estimation process 112f estimates the user's emotion by appropriately combining five results: the results from the language analysis process 112a, the results from the tone analysis process 112b, the results from the speed analysis process 112c, the results from the timing analysis process 112d, and the results from the facial expression analysis process 112e. Examples of the user's emotion that can be estimated include joy / anger / sadness / happiness, pleasure / displeasure, arousal / de-arousal, positive / negative, etc. Furthermore, in the facial expression analysis process 112e, the user's emotions are estimated in real time.

[0030] The evaluation unit 113 evaluates the user's exercise form based on the results of the analysis of the user's video (Vi) and the exercise instruction content (Eg) stored in the database (DB). The analysis process of the user's video (Vi) in the evaluation unit 113 is the posture estimation process 113a. The posture estimation process 113a focuses on the user's entire body in the user's video (Vi) and estimates the user's posture. For example, statically, the posture estimation process 113a focuses on the user's joints (shoulders, arms, wrists, ankles, etc.) and specific body parts (navel, etc.) and estimates the opening angles and coordinates of the joints and body parts. Dynamically, the posture estimation 113a estimates the amount of movement and direction of movement per unit time at the joints and body parts.

[0031] The exercise instruction content Eg includes exercise form data, exercise support commentary data, exercise menu data, and instruction policy data. Of these, the exercise form data is data that shows the posture and movements that the user should adopt during online lessons. In other words, the exercise form data is data that shows exemplary posture and movements. The exercise form data is supplied to the evaluation unit 113 as 113b.

[0032] The exercise form evaluation process 113c compares the user's exercise form estimated by the posture estimation process 113a with the exemplary exercise form shown in the exercise form data 113b to evaluate the user's exercise form in online lessons in real time. In evaluating exercise form, the evaluation results may include data indicating how closely the user's exercise form in the online lesson resembles the model exercise form. For example, in a given exercise, the model form might be one in which the straight line calculated by connecting the user's hand to the shoulder joint position remains horizontal for several seconds. The evaluation results could then be a binary value indicating whether the straight line calculated from the actual user's video is good if it is within 15 degrees of horizontal, or bad if it is not. Alternatively, the evaluation results could include an index (similarity) indicating how closely the opening angles and coordinates of the user's joints and body parts resemble the model.

[0033] Figure 6 shows the flow of various data and various processes in the instruction scenario generation unit 114. The instruction scenario generation unit 114 generates an instruction scenario based on the following data, information, and results. Specifically, the instruction scenario generation unit 114 uses emotion estimation results 114a, exercise form evaluation results 114b, exercise instruction data 114c, exercise menu data 114d, instruction judgment rule data 114e, user basic information 114f, dialogue data 114g, and coach information 114h when generating an instruction scenario.

[0034] The emotion estimation result 114a is the result of the estimation unit 112 estimating the emotions of the user taking the online lesson. The exercise form evaluation result 114b is the result of the evaluation unit 113 evaluating the exercise form of a user taking an online lesson.

[0035] Exercise instruction data 114c is a collection of data showing the coach's facial expressions, movements, and statements during online lessons. Although not specifically illustrated, data indicating the coach's facial expression can be obtained by focusing on the coach's face in the video Ci of the coach's entire body captured by the camera 150 and processing the coach's facial expression in the same way as the facial expression analysis process 112e. The coach's movements can be obtained by processing the video Ci of the coach's entire body captured by the camera 150 in the same way as the posture estimation process 113a. Data indicating the content of the coach's speech can be obtained by processing the coach's voice Co input by the voice input / output device 160 in the same way as the language analysis process 112a.

[0036] The exercise menu data 114d is data that shows the content of the lessons provided by the online lesson system 1, and is data that combines the type of lesson, the body part that is expected to be affected (loaded) by the lesson, and the intensity and difficulty level. Specifically, the exercise menu data 114d is data that combines the type of lesson, such as frailty prevention A, B, C, ..., yoga A, B, C, ..., dance A, B, C, ..., the body part, such as upper arms, ankles, whole body, and the intensity and difficulty level, such as soft / medium / hard. Coaching decision rule data 114e contains data that shows the criteria a coach uses when giving instructions, such as data indicating gentle / normal / strict. Furthermore, as mentioned above, the exercise menu data 114d and the instruction judgment rule data 114e are included in the exercise instruction content Eg.

[0037] User basic information 114f is information about the user receiving instruction from a coach in an online lesson, and includes information such as the user's name, age, gender, skill level, medical history, and exercise frequency. User basic information 114f is stored for each user in the user database of the storage device 120, for example, and in online lessons, the user basic information 114f corresponding to the user receiving instruction is read and used.

[0038] Dialogue data 114g is a collection of structured data that includes the flow and content of a dialogue. Dialogue data 114g is stored, for example, in a learning database within a database (DB).

[0039] Coach information 114h is information about the coach who instructs the user in the online lesson, and includes information such as the coach's age, gender, teaching experience, qualifications, areas of expertise, teaching philosophy, and life satisfaction. Coach information 114h is stored for each coach in the coach database of the storage device 120, for example, and in the online lesson, the coach information 114h corresponding to the court in which the user is being instructed is read and used.

[0040] The instruction scenario generation unit 114 generates an instruction scenario based on eight elements: emotion estimation results 114a, exercise form evaluation results 114b, exercise instruction data 114c, exercise menu data 114d, instruction judgment rule data 114e, user basic information 114f, dialogue data 114g, and coach information 114h. Furthermore, as will be described later, when generating instruction scenarios using machine learning, the instruction scenario generation unit 114 may estimate and output instruction scenarios using information on other coaches and users, as well as information on past instruction. Specifically, while the calculation methods for past satisfaction levels for similar users and instruction scenarios with high effectiveness are learned for each coach, the instruction scenario generation unit 114 generates instruction scenarios for users based on the results of this learning.

[0041] The instruction scenarios may be generated using only rule-based methods with 8 input elements, or using only machine learning methods with 8 input elements. Alternatively, some scenarios may be generated using machine learning, while the remaining scenarios are generated using rule-based methods. In machine learning, it is preferable to use training data based on correct answers from excellent coaches beforehand.

[0042] Next, we will explain how online lesson system 1 works. Online lessons are based on the premise that at least one user participates and that one coach instructs that user. Users who wish to participate in online lessons operate the terminal device 20 to access the coach support device 10 in advance and input the information listed in the user basic information 114f. The input information is stored in the user database of the database DB, linked to the user's identifier. Furthermore, the information listed in Coach Information 114h is assumed to already be stored in the coach database, linked to the coach's identifier.

[0043] Users who wish to participate in an online lesson must connect to the coach support device 10 by operating their terminal device 20 before the predetermined lesson start date and time, and notify the device of their intention to participate in the online lesson. This lesson will follow the content indicated in the exercise menu data 114d. When the lesson start time arrives, the control management unit 111 of the coach support device 10 identifies the users who have contacted them to participate before the start of the lesson.

[0044] Figure 7 is a flowchart showing the operation of the online lesson system 1 from the point in time when the users participating in the lesson are identified. The control management unit 111, having identified the users participating in the lesson, reads the basic user information corresponding to the identified users and transfers it to the instruction scenario generation unit 114. The control management unit 111 also transfers the coach information corresponding to the coach who will be instructing the lesson to the instruction scenario generation unit 114 (step Sa1).

[0045] Next, the control management unit 111 transfers the exercise form data 113b corresponding to the lesson from the exercise instruction content Eg stored in the database DB to the evaluation unit 113. The control management unit 111 also transfers the exercise menu data 114d and instruction judgment rule data 114e corresponding to the lesson from the exercise instruction content Eg to the instruction scenario generation unit 114 (step Sa2).

[0046] When an online lesson begins, the terminal device 20 captures the entire body of the user taking the lesson using the camera device 250, and the resulting video signal is transmitted to the coach support device 10. The user's voice is input to the audio input / output device, and the audio signal is transmitted to the coach support device 10. The control management unit 111 forwards the audio signal transmitted from the terminal device 20 as the user's voice Vo to the estimation unit 112, and forwards the video signal transmitted from the terminal device 20 as the user's video Vi to the estimation unit 112 and the evaluation unit 113.

[0047] The control management unit 111 determines whether or not it is time to generate an instruction scenario in the online lesson (step Sa3). The timing for generating a coaching scenario is the moment when the coach is expected to provide instruction to the user. Examples of timing for generating a coaching scenario include when the user speaks, when a certain amount of time has elapsed since the start of the lesson, when the user is estimated to be experiencing a specific emotion, when the user's exercise form evaluation result falls below a threshold, when the exercise form evaluation result rises above a threshold, or combinations of these timings. Furthermore, a user's exercise form evaluation result below the threshold indicates that the user's exercise form is inappropriate and should be corrected. Conversely, a user's exercise form evaluation result above the threshold indicates that the user's exercise form is appropriate and should be praised. Furthermore, the timing for generating the instruction scenario could also be at the beginning of the lesson, when the coach informs the user of the proposed exercise menu, if the menu is decided before the lesson. Alternatively, the scenario could be generated at the end of the lesson, when the coach reviews the lesson's results and provides feedback to the user, if there is a process of reviewing the lesson's outcomes and assigning homework.

[0048] If it is determined that it is not time to generate a guidance scenario (i.e., the result of step Sa3 is "No"), the control management unit 111 returns the processing procedure to step Sa3. Therefore, the processing procedure cycles through step Sa3 and enters a waiting state until it is determined that it is time to generate a guidance scenario.

[0049] If the control management unit 111 determines that it is time to generate a coaching scenario (if the result of the determination in step Sa3 is "Yes"), it instructs the coaching scenario generation unit 114 to generate a coaching scenario based on the eight elements. Subsequently, the control management unit 111 instructs the display unit 115 to display information based on the generated coaching scenario (step Sa4). Based on these instructions, coach advice and other information based on the coaching scenario are displayed in a predetermined area of ​​the display device 130.

[0050] After step Sa4, the control management unit 111 returns the processing procedure to step Sa3. Therefore, each time it is determined that it is time to generate a guidance scenario, a command is issued to generate a guidance scenario, and advice and other information based on the generated guidance scenario are displayed in a predetermined area of ​​the display device 130.

[0051] This process continues until the online lesson ends. Note that steps Sa3 and Sa4 will be executed for each user if there are multiple users taking the online lesson.

[0052] In the lesson, the coach will adopt exercise forms appropriate to the type and difficulty level of the lesson. Meanwhile, in the lesson, the user will adopt various exercise forms under the guidance of the coach. In the lesson, both the coach and the user will speak as appropriate.

[0053] Figure 8 shows an example of a screen displayed on the display device 130 of the coach support device 10 during an online lesson. This screen is an example where there are four users taking the lesson: A, B, C, and D. Figure 9 shows an example of a screen displayed on the display device 230 of the terminal device 20 operated by, for example, user A, one of the four users.

[0054] In the example shown in Figure 8, the display area of ​​the display device 130 is divided into a total of 6 squares: 2 vertically and 3 horizontally. For convenience, these 6 squares are distinguished as upper left, upper center, upper right, lower left, lower center, and lower right. The upper left square displays User A's full-body image Ua, User A's profile picture, and User A's name "AA". Similarly, the upper center, lower left, and lower center squares display User B's full-body image Ub, User C's full-body image Uc, and User D's full-body image Ud, in that order, along with their profile pictures and user names. The top right box displays the coach's full-body video (Ch) and profile picture. The bottom right box displays Ad1 and Ad2, advice based on the coaching scenario.

[0055] The black circles Bm superimposed on the whole-body images Ua, Ub, Uc, and Ud indicate the joints and specific body parts of users A, B, C, and D. These black circles are obtained by the posture estimation process 113a, which takes the user's video Vi as input. The black circles Bm serve as a guide when the coach evaluates the user's movement form.

[0056] In the example shown in Figure 9, the display device 230 shows scaled-down images of the coach's full-body video (Ch) and user A's own full-body video (Ua). The reason for displaying User A's full-body video Ua on User A's terminal device 20 is to allow User A to confirm how much their exercise form, as shown by their own full-body video Ua, differs from the exercise form shown by the coach's full-body video Ch.

[0057] Figure 9 shows an example of the display screen on user A's terminal device 20, but similarly, users B, C, and D will also see a reduced-size image of the coach's full body (Ch) and their own full body image.

[0058] Figure 8 shows an example where User B's exercise form evaluation result is below the threshold, and it is estimated that their emotions are not good. Specifically, this is an example where User B's joint angles in their arms and legs are slightly wider than those of the coach, and their facial expression appears distressed. In this case, it is time to generate a coaching scenario for User B (the judgment result in step Sa3 becomes "Yes"), so a coaching scenario for User B in this situation is generated.

[0059] Advice Ad1 is a message based on the instruction scenario generated at this time. To inform the coach that an instruction scenario for user B has been generated on the display device 130 screen, the presentation unit 115 displays a marker Bkb with a thick dashed border in the upper middle square corresponding to user B, and also makes the marker Bkm blink. Advice Ad1 is an example message based on the generated coaching scenario. By following the displayed Advice Ad1, even an inexperienced coach can instruct User B at the right time with the right content.

[0060] Furthermore, if the coach provides voice instructions to user B, the voice signal passes through the voice input / output device 160, communication device 140, network 30, communication device 240, and voice input / output device 260 in that order, and is converted into voice by the voice input / output device 260 and output. Furthermore, the coach may transmit the audio spoken during instruction only to user B's terminal device 20, or to include the terminal devices 20 of the other users A, C, and D as well.

[0061] Figure 8 also shows an example where the motor form evaluation results for users A and C are above the threshold and their emotions are estimated to be positive. Specifically, this is an example where the joint angles of the limbs of users A and C are not significantly different from those of the coach, and their facial expressions indicate a state of joy. In this case as well, a coaching scenario for users A and C is generated. Advice Ad2 is a message based on the instruction scenario generated at this time. To inform the coach that instruction scenarios for users A and C have been generated on the display device 130 screen, the presentation unit 115 displays and flashes a marker Bka with a thick solid border in the upper left square corresponding to user A, and a marker Bkc with a thick solid border in the lower left square corresponding to user C.

[0062] Advice Ad2 is an example message based on the generated coaching scenario. By following the displayed Advice Ad2, even an inexperienced coach can provide guidance that includes praising users A and C at the appropriate times.

[0063] The reason for using a dashed line for marker Bkb and solid lines for markers Bka and Bkc is that the coaching scenario for user B is based on negative results, while the coaching scenarios for users A and C are based on positive results. This difference is intended to intuitively inform the coach of the difference between the two. If this difference can be distinguished, it may be represented by line color rather than dashed or solid lines.

[0064] If User D speaks during the lesson, their statement may be transcribed and displayed in text Cv1, as shown in the middle-bottom square in Figure 8. In online lessons, the user and the terminal device 20 are separated because it is necessary to film the user's entire body, which may make it difficult to hear the user's voice. Displaying the statement as text in this way prevents the coach from missing what the user says.

[0065] The coach can evaluate the full-body video (Ua, Ub, Uc) and provide advice that differs from Ad1 and Ad2, or they can choose not to provide any guidance if they deem it unnecessary. In other words, the decision of whether or not to guide the user according to the guidance scenario is left to the coach's own judgment. In online lessons, the video and audio of both the user and the coach are recorded. This recording can be analyzed to provide coaches with feedback on what went well, what could have been improved, and whether the timing of their instruction was appropriate. Additionally, if a change in the coach's skill level is recognized after the online lesson has ended, the coach's information may be updated.

[0066] Thus, according to the coach support device 10 of this embodiment, it is possible to support coaches so that they can provide appropriate instruction to users in online lessons at the appropriate time.

[0067] <Variation> In the embodiments described above, the full-body video of the coach (Ch) was displayed on the display devices 130 and 230. However, instead of the coach, the display unit 115 may display an avatar representing the coach. Figure 10 shows an example of a screen in which the display unit 115 displays an avatar At representing the coach in place of the full-body video of the coach (Ch) in the upper right square of the display device 130. Figure 11 shows an example of a screen in which the avatar At is displayed on the display device 230 of the terminal device 20 operated by user A.

[0068] The avatar At may be given movements similar to those of the coach's full-body video channel. More specifically, the coach's posture may be estimated from the coach's full-body video channel, the opening angles and coordinates of joints and body parts may be determined, and the parts corresponding to the joints and body parts in the avatar At may be given movements corresponding to the opening angles and coordinates. Furthermore, Avatar At may be given movements that follow the exemplary exercise form data 113b.

[0069] Even if an avatar (At) is displayed instead of the coach's full-body video (Ch), from the perspective of coach support, it is preferable to have a configuration where the coach is real, a coaching scenario is generated from eight elements, advice based on that coaching scenario is displayed, and the coach provides instruction based on that advice. In this configuration, data characterizing the coach's voice quality may be pre-stored in the database DB of the storage device 120, and the presentation unit 115 may synthesize a message based on the generated coaching scenario and have the avatar At speak in place of the coach. Furthermore, when one coach is giving online lessons to multiple users, if one user is having a high-priority conversation with the coach, and another user has a question, the presentation unit 115 may have a function to assist the coach, such as having the avatar At speak to the other user on behalf of the coach.

[0070] <Other> The block diagram used in the description of the above embodiment shows functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Furthermore, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one device that is physically or logically coupled, or it may be realized using two or more physically or logically separated devices that are directly or indirectly connected (for example, using wired, wireless, etc.). A functional block may be realized by combining the above one device or the above multiple devices with software. Functions include, but are not limited to, judgment, decision, determination, calculation, calculation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, assumption, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating (mapping), and assigning. For example, the functional block (component) that enables transmission is called a transmitting unit or transmitter. As mentioned above, the method of implementation is not particularly limited.

[0071] Information notification is not limited to the embodiments described herein and may be carried out by other means. For example, information notification may be carried out by physical layer signaling (e.g., DCI (Downlink Control Information), UCI (Uplink Control Information)), upper layer signaling (e.g., RRC (Radio Resource Control) signaling, MAC (Medium Access Control) signaling, broadcast information (MIB (Master Information Block), SIB (System Information Block))), other signals, or combinations thereof. RRC signaling may also be called RRC messages, and may be, for example, RRC Connection Setup messages, RRC Connection Reconfiguration messages, etc.

[0072] Each aspect / embodiment described herein may apply to systems utilizing LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G (4th generation mobile communication system), 5G (5th generation mobile communication system), 6th generation mobile communication system (6G), xth generation mobile communication system (xG) (xG (where x is, for example, an integer or decimal)), FRA (Future Radio Access), NR (new Radio), New radio access (NX), Future generation radio access (FX), W-CDMA®, GSM®, CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi®), IEEE 802.16 (WiMAX®), IEEE 802.20, UWB (Ultra-Wide Band), Bluetooth®, and other appropriate systems, as well as at least one of the next-generation systems that are extended, modified, created, or defined based thereon. Furthermore, multiple systems may be applied in combination (for example, a combination of at least one of LTE and LTE-A with 5G).

[0073] The processing procedures, sequences, flowcharts, etc., of each aspect / embodiment described herein may be reordered, provided they are consistent with each other. For example, the methods described herein present various step elements in an exemplary order and are not limited to that specific order.

[0074] In this disclosure, certain operations performed by a base station may, in some cases, be performed by its upper node. In a network consisting of one or more network nodes having a base station, it is clear that various operations performed for communication with a terminal can be performed by the base station and at least one other network node (for example, an MME or S-GW, but not limited to these). Although the above example illustrates the case where there is one other network node besides the base station, it may also be a combination of multiple other network nodes (for example, an MME and an S-GW).

[0075] Information (described later) can be output from a higher layer (or lower layer) to a lower layer (or higher layer). Input and output may also occur via multiple network nodes.

[0076] Input and output information may be stored in a specific location (e.g., memory) or managed using a management table. Input and output information may be overwritten, updated, or appended to. Output information may be deleted. Input information may be transmitted to other devices.

[0077] The determination may be made by a value represented by 1 bit (0 or 1), by a boolean value (true or false), or by a numerical comparison (for example, a comparison with a predetermined value).

[0078] Each aspect / embodiment described herein may be used individually, in combination, or switched between as needed during implementation. Furthermore, notification of specific information (e.g., notification that "X is") is not limited to explicit notification, but may also be implicit (e.g., by not providing such notification).

[0079] Although the present disclosure has been described in detail above, it will be clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the intent and scope of the present disclosure as defined by the claims. Accordingly, the descriptions in the present disclosure are illustrative and not intended to be restrictive in any way.

[0080] Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, and so on, whether they are called software, firmware, middleware, microcode, hardware description languages, or by any other name. Furthermore, software, instructions, information, etc., may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using at least one of wired technologies (such as coaxial cable, fiber optic cable, twisted pair, or digital subscriber line (DSL)) and wireless technologies (such as infrared or microwave), then at least one of these wired and wireless technologies is included in the definition of a transmission medium.

[0081] The information, signals, etc., described in this disclosure may be represented using any of the following different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc., which may be referred to throughout the above description, may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof. Terms used in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). A signal may also be a message. Furthermore, a component carrier (CC) may also be called a carrier frequency, cell, frequency carrier, etc.

[0082] The terms “system” and “network” as used in this disclosure are interchangeable.

[0083] Furthermore, the information, parameters, etc., described in this disclosure may be expressed using absolute values, relative values ​​from a given value, or other corresponding information. For example, radio resources may be indicated by an index. The names used for the parameters described above are not limiting in any way. Moreover, the formulas, etc., that use these parameters may differ from those expressly disclosed in this disclosure. Various channels (e.g., PUCCH, PDCCH, etc.) and information elements can be identified by any suitable name, and therefore, the various names assigned to these various channels and information elements are not limiting in any way.

[0084] In this disclosure, terms such as “Base Station (BS),” “wireless base station,” “fixed station,” “NodeB,” “eNodeB (eNB),” “gNodeB (gNB),” “access point,” “transmission point,” “reception point,” “transmission / reception point,” “cell,” “sector,” “cell group,” “carrier,” and “component carrier” may be used interchangeably. A base station may also be referred to by terms such as macrocell, small cell, femtocell, and picocell. A base station may house one or more (e.g., three) cells. When a base station accommodates multiple cells, the entire coverage area of ​​the base station can be divided into multiple smaller areas, each of which may be provided with communication services by a base station subsystem (for example, a Remote Radio Head (RRH)). The terms “cell” or “sector” refer to part or all of the coverage area of ​​at least one of the base station and / or base station subsystems that provide communication services in this coverage. In this disclosure, the transmission of information by a base station to a terminal may be interpreted as the base station instructing the terminal to perform information-based control or operation.

[0085] In this disclosure, terms such as “Mobile Station (MS),” “user terminal,” “User Equipment (UE),” and “terminal” may be used interchangeably. A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or several other appropriate terms.

[0086] . At least one of the base station and the mobile station may be called a transmitting device, a receiving device, a communication device, etc. At least one of the base station and the mobile station may also be a device mounted on a mobile body, the mobile body itself, etc. The mobile body refers to a movable object, and its speed of movement is arbitrary. This also includes the case when the mobile body is stationary. The mobile body includes, but is not limited to, vehicles, transport vehicles, automobiles, motorcycles, bicycles, connected cars, excavators, bulldozers, wheel loaders, dump trucks, forklifts, trains, buses, handcarts, rickshaws, ships and other watercraft, airplanes, rockets, satellites, drones (registered trademark), multicopters, quadcopters, balloons, and items mounted on them. The mobile body may also be a mobile body that moves autonomously based on operation commands. It may be a vehicle (e.g., a car, an airplane, etc.), an unmanned mobile body (e.g., a drone, an autonomous vehicle, etc.), or a robot (manned or unmanned). Furthermore, at least one of the base station and the mobile station may include devices that do not necessarily move during communication operations. For example, at least one of the base station and the mobile station may be an IoT (Internet of Things) device such as a sensor. Also, the term "base station" in this disclosure may be interpreted as "user terminal." For example, each aspect / embodiment of this disclosure may be applied to a configuration in which communication between a base station and a user terminal is replaced with communication between multiple user terminals (which may be called, for example, D2D (Device-to-Device), V2X (Vehicle-to-Everything)). In this case, the user terminal may have the functions that the base station has. Also, terms such as "uplink" and "downlink" may be interpreted as terms corresponding to terminal-to-terminal communication (for example, "side"). For example, uplink channel, downlink channel, etc. may be interpreted as side channel. Similarly, the term "user terminal" in this disclosure may be interpreted as "base station." In this case, the base station may have the functions that the user terminal has.

[0087] As used in this disclosure, the terms “determining” and “determining” may encompass a wide variety of actions. “Determining” may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, or inquiring (e.g., searching in a table, database, or other data structure), or ascertaining. “Determining” may also include, for example, receiving (e.g., receiving information), transmitting (e.g., sending information), inputting, outputting, or accessing (e.g., accessing data in memory). Furthermore, "judgment" and "decision" can include considering something as having been "judged" or "decided" after resolving, selecting, choosing, establishing, comparing, etc. In other words, "judgment" and "decision" can include considering something as having been "judged" or "decided" after some action. Also, "judgment (decision)" can be reinterpreted as "assuming," "expecting," or "considering."

[0088] The terms “connected,” “coupled,” or any variation thereof, mean any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are “connected” or “coupled” with each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, “connection” may be reinterpreted as “access.” As used in this disclosure, two elements may be considered to be “connected” or “coupled” with each other using at least one of one or more wires, cables, and printed electrical connections, and, in some non-limiting and non-exclusive examples, electromagnetic energy having wavelengths in the radio frequency domain, microwave domain, and optical (both visible and invisible) domain.

[0089] The reference signal can also be abbreviated as RS (Reference Signal), and may be called a pilot depending on the applicable standard.

[0090] In this disclosure, the phrase "based on" does not mean "based solely on" unless otherwise specified. In other words, the phrase "based on" means both "based solely on" and "based at least on."

[0091] Any reference to elements using designations such as “first,” “second,” etc., as used in this disclosure does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient way to distinguish between two or more elements. Accordingly, references to first and second elements do not imply that only two elements may be adopted, or that the first element must precede the second element in any way.

[0092] In the configuration of each of the above devices, "means" may be replaced with "part," "circuit," "device," etc.

[0093] Where the terms “include,” “including,” and their variations are used in this disclosure, these terms are intended to be inclusive, as is the term “comprising.” Furthermore, the term “or” as used in this disclosure is not intended to mean exclusive OR.

[0094] In this disclosure, if articles are added by translation, such as a, an, and the in English, this disclosure may include the fact that the noun following these articles is plural.

[0095] In this disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "combine" may be interpreted similarly to "different."

[0096] It will be apparent to those skilled in the art that the present invention is not limited to the embodiments described herein. The present invention can be implemented in modified and altered forms without departing from the spirit and scope of the invention as defined by the claims. Accordingly, the description herein is for illustrative purposes only and is not intended to be restrictive in any way to the present invention. Furthermore, multiple embodiments selected from those illustrated herein may be combined. [Explanation of symbols]

[0097] 1...Online lesson system, 10...Coach support device, 20...Terminal device, 30...Network, 110...Processing device, 111...Control management unit, 112...Estimation unit, 113...Evaluation unit, 114...Instruction scenario generation unit, 114a...Emotion estimation result, 114b...Exercise form evaluation result, 114c...Exercise instruction data, 114d...Exercise menu data, 114e...Instruction judgment rule data, 114f...User basic information, 114g...Dialogue data, 114h...Coach basic information, 115...Presentation unit.

Claims

1. A coach support device that assists a coach in an online lesson, which receives a video signal and an audio signal from a user receiving a lesson from a coach via a network from a terminal device, An estimation unit that estimates the user's emotions based on the video signal and the audio signal, An evaluation unit that evaluates the user's exercise form based on the aforementioned video signal, At a minimum, a generation unit generates a guidance scenario for the user based on the emotion estimation result by the estimation unit and the exercise form evaluation result by the evaluation unit, A presentation unit that presents a message based on the aforementioned coaching scenario to the coach, It has, When multiple users take a lesson at the same time, The aforementioned display unit is, Provide information indicating which users the proposed instructional scenario is intended for. A coaching support device for online lessons.

2. The estimation unit, Based on at least one of the user's facial expression shown in the video signal and the user's statement shown in the audio signal, the user's emotions are estimated. The evaluation unit described above, The user's exercise form is evaluated based on a comparison between the user's exercise form shown in the aforementioned video signal and an exemplary exercise form. The coach support device for online lessons according to claim 1.

3. The display portion is, The coach is presented with a message based on the aforementioned coaching scenario. The coach support device for online lessons according to claim 1.

4. The display portion is, The system presents the coach with a message based on the aforementioned instructional scenario and causes the coach to output an audio signal to the terminal device. The coach support device for online lessons according to claim 1.

5. The aforementioned instructional scenario generation unit, moreover Exercise instruction data showing the content of the aforementioned coach's instruction, Exercise menu data showing the content of the lesson, Instructional decision-making rule data that shows the teaching policy for the lesson, Basic user information identifying the user, Dialogue data showing the content of the conversation the coach should have, Information about the coach who instructs the user, Based on this, generate a training scenario for the user. The coach support device for online lessons according to claim 1.

Citation Information

Patent Citations

  • Trainer support system, trainer support method and program

    JP2002346012A

  • Exercise instruction system and its management device

    JP2006259929A

  • Exercise support system, user terminal therefor and exercise support program

    JP2006302122A

  • Video generating apparatus and method, and program

    JP2013111449A

  • Treatment and / or exercise guidance process management system, and program, computer device, and method for treatment and / or exercise guidance process management

    JP2021129992A